{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,9]],"date-time":"2026-05-09T17:21:35Z","timestamp":1778347295798,"version":"3.51.4"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,6,16]],"date-time":"2023-06-16T00:00:00Z","timestamp":1686873600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,6,30]]},"abstract":"<jats:p>The multi-modal machine translation task is to improve translation quality with the help of additional visual input. It is expected to disambiguate or complement semantics while there are ambiguous words or incomplete expressions in the sentences. Existing methods have tried many ways to fuse visual information into text representations. However, only a minority of sentences need extra visual information as complementary. Without guidance, models tend to learn text-only translation from the major well-aligned translation pairs. In this article, we propose a contrastive adversarial training approach to enhance visual participation in semantic representation learning. By contrasting multi-modal input with the adversarial samples, the model learns to identify the most informed sample that is coupled with a congruent image and several visual objects extracted from it. This approach can prevent the visual information from being ignored and further fuse cross-modal information. We examine our method in three multi-modal language pairs. Experimental results show that our model is capable of improving translation accuracy. Further analysis shows that our model is more sensitive to visual information.<\/jats:p>","DOI":"10.1145\/3587267","type":"journal-article","created":{"date-parts":[[2023,3,14]],"date-time":"2023-03-14T12:08:46Z","timestamp":1678795726000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Contrastive Adversarial Training for Multi-Modal Machine Translation"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7327-377X","authenticated-orcid":false,"given":"Xin","family":"Huang","sequence":"first","affiliation":[{"name":"University of Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5293-7434","authenticated-orcid":false,"given":"Jiajun","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1977-1115","authenticated-orcid":false,"given":"Chengqing","family":"Zong","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,6,16]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations (ICLR\u201915)","author":"Bahdanau Dzmitry","year":"2015","unstructured":"Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In Proceedings of the 3rd International Conference on Learning Representations (ICLR\u201915). http:\/\/arxiv.org\/abs\/1409.0473."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-6402"},{"key":"e_1_3_2_4_2","volume-title":"Proceedings of the 6th International Conference on Learning Representations (ICLR\u201918)","author":"Belinkov Yonatan","year":"2018","unstructured":"Yonatan Belinkov and Yonatan Bisk. 2018. Synthetic and natural noise both break neural machine translation. In Proceedings of the 6th International Conference on Learning Representations (ICLR\u201918). https:\/\/openreview.net\/forum?id=BJ8vJebC-."},{"key":"e_1_3_2_5_2","article-title":"Multimodal attention for neural machine translation","volume":"1609","author":"Caglayan Ozan","year":"2016","unstructured":"Ozan Caglayan, Lo\u00efc Barrault, and Fethi Bougares. 2016. Multimodal attention for neural machine translation. CoRR abs\/1609.03976 (2016).","journal-title":"CoRR"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1422"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1105"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1175"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1425"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-3348"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1329"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-4718"},{"key":"e_1_3_2_13_2","unstructured":"Desmond Elliott Stella Frank and Eva Hasler. 2015. Multilingual image description with neural sequence models. (2015). arXiv:cs.CL\/1510.04709 (2015)."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-3210"},{"key":"e_1_3_2_15_2","first-page":"130","volume-title":"Proceedings of the 8th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Elliott Desmond","year":"2017","unstructured":"Desmond Elliott and \u00c1kos K\u00e1d\u00e1r. 2017. Imagination improves multimodal translation. In Proceedings of the 8th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 130\u2013141. https:\/\/aclanthology.org\/I17-1014."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.390"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_18_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919)","author":"Hjelm R. Devon","year":"2019","unstructured":"R. Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Philip Bachman, Adam Trischler, and Yoshua Bengio. 2019. Learning deep representations by mutual information estimation and maximization. In Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919). https:\/\/openreview.net\/forum?id=Bklr3j0cKX."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.repl4nlp-1.30"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-2360"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-emnlp.92"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.439"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1653"},{"key":"e_1_3_2_24_2","article-title":"A survey on contrastive self-supervised learning","volume":"2011","author":"Jaiswal Ashish","year":"2020","unstructured":"Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh, Debapriya Banerjee, and Fillia Makedon. 2020. A survey on contrastive self-supervised learning. CoRR abs\/2011.00362 (2020).","journal-title":"CoRR"},{"key":"e_1_3_2_25_2","volume-title":"Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (NeurIPS\u201920)","author":"Khosla Prannay","year":"2020","unstructured":"Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (NeurIPS\u201920). https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/d89a66c7c80a29b1bdbab0f2a1a94af8-Abstract.html."},{"key":"e_1_3_2_26_2","volume-title":"Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (NeurIPS\u201920)","author":"Kim Minseon","year":"2020","unstructured":"Minseon Kim, Jihoon Tack, and Sung Ju Hwang. 2020. Adversarial self-supervised contrastive learning. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (NeurIPS\u201920). https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/1f1baa5b8edac74eb4eaa329f14a0361-Abstract.html."},{"key":"e_1_3_2_27_2","unstructured":"Wonjae Kim Bokyung Son and Ildoo Kim. 2021. ViLT: Vision-and-language transformer without convolution or region supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML\u201921) . 5583\u20135594. http:\/\/proceedings.mlr.press\/v139\/kim21k.html."},{"key":"e_1_3_2_28_2","article-title":"Unifying visual-semantic embeddings with multimodal neural language models","volume":"1411","author":"Kiros Ryan","year":"2014","unstructured":"Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2014. Unifying visual-semantic embeddings with multimodal neural language models. CoRR abs\/1411.2539 (2014).","journal-title":"CoRR"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.671"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.438"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.673"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1230"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.202"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-2031"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-6326"},{"key":"e_1_3_2_37_2","article-title":"Gumbel-attention for multi-modal machine translation","volume":"2103","author":"Liu Pengbo","year":"2021","unstructured":"Pengbo Liu, Hailong Cao, and Tiejun Zhao. 2021. Gumbel-attention for multi-modal machine translation. CoRR abs\/2103.08862 (2021).","journal-title":"CoRR"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1166"},{"key":"e_1_3_2_39_2","first-page":"4204","volume-title":"Proceedings of the 12th Language Resources and Evaluation Conference","author":"Nakayama Hideki","year":"2020","unstructured":"Hideki Nakayama, Akihiro Tamura, and Takashi Ninomiya. 2020. A visually-grounded parallel corpus with phrase-to-region linking. In Proceedings of the 12th Language Resources and Evaluation Conference. 4204\u20134210. https:\/\/aclanthology.org\/2020.lrec-1.518."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.21"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.13053\/cys-23-4-3294"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0965-7"},{"key":"e_1_3_2_44_2","volume-title":"Proceedings of the 8th International Conference on Learning Representations (ICLR\u201920)","author":"Su Weijie","year":"2020","unstructured":"Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. VL-BERT: Pre-training of generic visual-linguistic representations. In Proceedings of the 8th International Conference on Learning Representations (ICLR\u201920). https:\/\/openreview.net\/forum?id=SygXPaEYvH."},{"key":"e_1_3_2_45_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems (NeurIPS\u201917)","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems (NeurIPS\u201917). 5998\u20136008. https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.181"},{"key":"e_1_3_2_47_2","first-page":"2720","volume-title":"Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI\u201921), the 33rd Conference on Innovative Applications of Artificial Intelligence (IAAI\u201921), and the 11th Symposium on Educational Advances in Artificial Intelligence (EAAI\u201921)","author":"Wang Dexin","year":"2021","unstructured":"Dexin Wang and Deyi Xiong. 2021. Efficient object-level visual context modeling for multimodal machine translation: Masking irrelevant objects helps grounding. In Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI\u201921), the 33rd Conference on Innovative Applications of Artificial Intelligence (IAAI\u201921), and the 11th Symposium on Educational Advances in Artificial Intelligence (EAAI\u201921). 2720\u20132728. https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/view\/16376."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.480"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3457571"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00478"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.400"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.273"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00166"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1400"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3587267","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3587267","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:36:38Z","timestamp":1750282598000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3587267"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,16]]},"references-count":53,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,6,30]]}},"alternative-id":["10.1145\/3587267"],"URL":"https:\/\/doi.org\/10.1145\/3587267","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,16]]},"assertion":[{"value":"2022-09-20","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-02","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}