{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T18:41:03Z","timestamp":1784227263400,"version":"3.55.0"},"reference-count":227,"publisher":"Springer Science and Business Media LLC","issue":"2-3","license":[{"start":{"date-parts":[[2020,8,13]],"date-time":"2020-08-13T00:00:00Z","timestamp":1597276800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,8,13]],"date-time":"2020-08-13T00:00:00Z","timestamp":1597276800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100007601","name":"Horizon 2020","doi-asserted-by":"publisher","award":["780069"],"award-info":[{"award-number":["780069"]}],"id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007601","name":"Horizon 2020","doi-asserted-by":"publisher","award":["771113"],"award-info":[{"award-number":["771113"]}],"id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007601","name":"Horizon 2020","doi-asserted-by":"publisher","award":["678017"],"award-info":[{"award-number":["678017"]}],"id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100010897","name":"Newton Fund","doi-asserted-by":"publisher","award":["352343575"],"award-info":[{"award-number":["352343575"]}],"id":[{"id":"10.13039\/100010897","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Machine Translation"],"published-print":{"date-parts":[[2020,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Multimodal machine translation involves drawing information from more than one modality, based on the assumption that the additional modalities will contain useful alternative views of the input data. The most prominent tasks in this area are spoken language translation, image-guided translation, and video-guided translation, which exploit audio and visual modalities, respectively. These tasks are distinguished from their monolingual counterparts of speech recognition, image captioning, and video captioning by the requirement of models to generate outputs in a different language. This survey reviews the major data resources for these tasks, the evaluation campaigns concentrated around them, the state of the art in end-to-end and pipeline approaches, and also the challenges in performance evaluation. The paper concludes with a discussion of directions for future research in these areas: the need for more expansive and challenging datasets, for targeted evaluations of model performance, and for multimodality in both the input and output space.<\/jats:p>","DOI":"10.1007\/s10590-020-09250-0","type":"journal-article","created":{"date-parts":[[2020,8,13]],"date-time":"2020-08-13T12:02:36Z","timestamp":1597320156000},"page":"97-147","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":53,"title":["Multimodal machine translation through visuals and speech"],"prefix":"10.1007","volume":"34","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4929-4850","authenticated-orcid":false,"given":"Umut","family":"Sulubacak","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ozan","family":"Caglayan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stig-Arne","family":"Gr\u00f6nroos","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Aku","family":"Rouhe","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Desmond","family":"Elliott","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lucia","family":"Specia","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"J\u00f6rg","family":"Tiedemann","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,8,13]]},"reference":[{"key":"9250_CR1","unstructured":"Abdelali A, Guzman F, Sajjad H, Vogel S (2014) The AMARA corpus: building parallel language resources for the educational domain. In: Proceedings of the 9th international conference on language resources and evaluation (LREC), European Language Resources Association (ELRA), Reykjav\u00edk, Iceland, pp 1856\u20131862"},{"key":"9250_CR2","unstructured":"Akiba Y, Federico M, Kando N, Nakaiwa H, Paul M, Tsujii J (2004) Overview of the IWSLT 2004 evaluation campaign. In: Proceedings of the 2004 international workshop on spoken language translation (IWSLT), Kyoto, pp 1\u201312"},{"key":"9250_CR3","doi-asserted-by":"crossref","unstructured":"Anastasopoulos A, Chiang D (2018) Tied multitask learning for neural speech translation. In: Proceedings of the 2018 conference of the North American chapter of the association for computational linguistics: human language technologies (NAACL-HLT), Association for Computational Linguistics (ACL), New Orleans, Louisiana, pp 82\u201391","DOI":"10.18653\/v1\/N18-1008"},{"key":"9250_CR4","doi-asserted-by":"crossref","unstructured":"Anastasopoulos A, Chiang D, Duong L (2016) An unsupervised probability model for speech-to-translation alignment of low-resource languages. In: Proceedings of the 2016 conference on empirical methods in natural language processing (EMNLP), Association for Computational Linguistics (ACL), Austin, pp 1255\u20131263","DOI":"10.18653\/v1\/D16-1133"},{"issue":"2","key":"9250_CR5","doi-asserted-by":"publisher","first-page":"356","DOI":"10.1109\/TASL.2011.2125954","volume":"20","author":"X Anguera","year":"2012","unstructured":"Anguera X, Bozonnet S, Evans N, Fredouille C, Friedland G, Vinyals O (2012) Speaker diarization: a review of recent research. IEEE Trans Audio Speech Lang Process 20(2):356\u2013370","journal-title":"IEEE Trans Audio Speech Lang Process"},{"key":"9250_CR6","doi-asserted-by":"crossref","unstructured":"Antol S, Agrawal A, Lu J, Mitchell M, Batra D, Lawrence\u00a0Zitnick C, Parikh D (2015) VQA: visual question answering. In: Proceedings of the 2015 IEEE international conference on computer vision (ICCV), Santiago, pp 2425\u20132433","DOI":"10.1109\/ICCV.2015.279"},{"key":"9250_CR7","unstructured":"Arslan HS, Fishel M, Anbarjafari G (2018) Doubly attentive transformer machine translation. Computing research repository. arXiv:1807.11605"},{"key":"9250_CR8","unstructured":"Bahar P, Zeyer A, Schl\u00fcter R, Ney H (2019) On using SpecAugment for end-to-end speech translation. In: Proceedings of the 16th international workshop on spoken language translation (IWSLT), Hong Kong"},{"key":"9250_CR9","unstructured":"Bahdanau D, Cho K, Bengio Y (2015) Neural machine translation by jointly learning to align and translate. In: Proceedings of the 3rd international conference on learning representations (ICLR), San Diego"},{"key":"9250_CR10","unstructured":"Baltru\u0161aitis T, Ahuja C, Morency LP (2017) Multimodal machine learning: a survey and taxonomy. Computing research repository arXiv:1705.09406"},{"key":"9250_CR11","doi-asserted-by":"crossref","unstructured":"Bansal S, Kamper H, Lopez A, Goldwater S (2017) Towards speech-to-text translation without speech recognition. In: Proceedings of the 15th conference of the european chapter of the association for computational linguistics (EACL), Association for computational linguistics (ACL), Valencia, pp 474\u2013479","DOI":"10.18653\/v1\/E17-2076"},{"key":"9250_CR12","doi-asserted-by":"crossref","unstructured":"Bansal S, Kamper H, Livescu K, Lopez A, Goldwater S (2018) Low-resource speech-to-text translation. In: Proceedings of interspeech, Hyderabad, pp 1298\u20131302","DOI":"10.21437\/Interspeech.2018-1326"},{"key":"9250_CR13","doi-asserted-by":"crossref","unstructured":"Bansal S, Kamper H, Livescu K, Lopez A, Goldwater S (2019) Pre-training on high-resource speech recognition improves low-resource speech-to-text translation. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies (NAACL-HLT), Association for Computational Linguistics (ACL), Minneapolis, pp 58\u201368","DOI":"10.18653\/v1\/N19-1006"},{"key":"9250_CR14","doi-asserted-by":"crossref","unstructured":"Barrault L, Bougares F, Specia L, Lala C, Elliott D, Frank S (2018) Findings of the third shared task on multimodal machine translation. In: Proceedings of the 3rd conference on machine translation (WMT), association for computational linguistics (ACL), Belgium, pp 308\u2013327","DOI":"10.18653\/v1\/W18-6402"},{"key":"9250_CR15","unstructured":"Belinkov Y, Bisk Y (2018) Synthetic and natural noise both break neural machine translation. In: Proceedings of the 6th international conference on learning representations (ICLR), Vancouver"},{"issue":"Feb","key":"9250_CR16","first-page":"1137","volume":"3","author":"Y Bengio","year":"2003","unstructured":"Bengio Y, Ducharme R, Vincent P, Jauvin C (2003) A neural probabilistic language model. J Mach Learn Res 3(Feb):1137\u20131155","journal-title":"J Mach Learn Res"},{"key":"9250_CR17","unstructured":"Bentivogli L, Cettolo M, Federico M, Federmann C (2018) Machine translation human evaluation: an investigation of evaluation based on post-editing and its relation with direct assessment. In: Proceedings of the 2018 international workshop on spoken language translation (IWSLT), Bruges, pp 62\u201369"},{"key":"9250_CR18","unstructured":"B\u00e9rard A, Pietquin O, Servan C, Besacier L (2016) Listen and translate: A proof of concept for end-to-end speech-to-text translation. In: Proceedings of the 29th neural information processing systems conference (NeurIPS) end-to-end learning for speech and audio processing workshop, Barcelona"},{"key":"9250_CR19","doi-asserted-by":"crossref","unstructured":"B\u00e9rard A, Besacier L, Kocabiyikoglu AC, Pietquin O (2018) End-to-end automatic speech translation of audiobooks. 2018 international conference on acoustics, speech and signal processing (ICASSP). IEEE, Calgary, pp 6224\u20136228","DOI":"10.1109\/ICASSP.2018.8461690"},{"key":"9250_CR20","doi-asserted-by":"publisher","first-page":"409","DOI":"10.1613\/jair.4900","volume":"55","author":"R Bernardi","year":"2016","unstructured":"Bernardi R, Cakici R, Elliott D, Erdem A, Erdem E, Ikizler-Cinbis N, Keller F, Muscat A, Plank B (2016) Automatic description generation from images: a survey of models, datasets, and evaluation measures. J Artif Intell Res 55:409\u2013442","journal-title":"J Artif Intell Res"},{"key":"9250_CR21","unstructured":"Boito MZ, Havard WN, Garnerin M, Ferrand \u00c9L, Besacier L (2019) MaSS: a large and clean multilingual corpus of sentence-aligned spoken utterances extracted from the Bible. Computing research repository arXiv:1907.12895"},{"key":"9250_CR22","unstructured":"Caglayan O (2019) Multimodal machine translation. PhD thesis, Universit\u00e9 du Maine"},{"key":"9250_CR23","doi-asserted-by":"crossref","unstructured":"Caglayan O, Aransa W, Wang Y, Masana M, Garc\u00eda-Mart\u00ednez M, Bougares F, Barrault L, van\u00a0de Weijer J (2016a) Does multimodality help human and machine for translation and image captioning? In: Proceedings of the 1st conference on machine translation (WMT), association for computational linguistics (ACL), Berlin, pp 627\u2013633","DOI":"10.18653\/v1\/W16-2358"},{"key":"9250_CR24","unstructured":"Caglayan O, Barrault L, Bougares F (2016b) Multimodal attention for neural machine translation. Computing research repository arXiv:1609.03976"},{"key":"9250_CR25","doi-asserted-by":"crossref","unstructured":"Caglayan O, Aransa W, Bardet A, Garc\u00eda-Mart\u00ednez M, Bougares F, Barrault L, Masana M, Herranz L, van\u00a0de Weijer J (2017a) LIUM-CVC submissions for WMT17 multimodal translation task. In: Proceedings of the 2nd conference on machine translation, association for computational linguistics (ACL), Copenhagen, pp 432\u2013439","DOI":"10.18653\/v1\/W17-4746"},{"key":"9250_CR26","doi-asserted-by":"publisher","first-page":"15","DOI":"10.1515\/pralin-2017-0035","volume":"109","author":"O Caglayan","year":"2017","unstructured":"Caglayan O, Garc\u00eda-Mart\u00ednez M, Bardet A, Aransa W, Bougares F, Barrault L (2017b) NMTPY: a flexible toolkit for advanced neural machine translation systems. Prague Bull Math Linguistics 109:15\u201328","journal-title":"Prague Bull Math Linguistics"},{"key":"9250_CR27","doi-asserted-by":"crossref","unstructured":"Caglayan O, Bardet A, Bougares F, Barrault L, Wang K, Masana M, Herranz L, van\u00a0de Weijer J (2018) LIUM-CVC submissions for WMT18 multimodal translation task. In: Proceedings of the 3rd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Belgium, pp 603\u2013608","DOI":"10.18653\/v1\/W18-6438"},{"key":"9250_CR28","doi-asserted-by":"crossref","unstructured":"Caglayan O, Madhyastha P, Specia L, Barrault L (2019) Probing the need for visual context in multimodal machine translation. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies (NAACL-HLT), Association for Computational Linguistics (ACL), Minneapolis, pp 4159\u20134170","DOI":"10.18653\/v1\/N19-1422"},{"key":"9250_CR29","doi-asserted-by":"crossref","unstructured":"Calixto I, Liu Q (2017a) Incorporating global visual features into attention-based neural machine translation. In: Proceedings of the 2017 conference on empirical methods in natural language processing (EMNLP), Association for Computational Linguistics (ACL), Copenhagen, pp 992\u20131003","DOI":"10.18653\/v1\/D17-1105"},{"key":"9250_CR30","doi-asserted-by":"crossref","unstructured":"Calixto I, Liu Q (2017b) Sentence-level multilingual multi-modal embedding for natural language processing. In: Proceedings of the international conference recent advances in natural language processing (RANLP), INCOMA Ltd., Varna, Bulgaria, pp 139\u2013148","DOI":"10.26615\/978-954-452-049-6_020"},{"key":"9250_CR31","doi-asserted-by":"crossref","unstructured":"Calixto I, Elliott D, Frank S (2016) Dcu-uva multimodal mt system report. In: Proceedings of the 1st conference on machine translation (WMT), Association for Computational Linguistics (ACL), Berlin, pp 634\u2013638","DOI":"10.18653\/v1\/W16-2359"},{"key":"9250_CR32","doi-asserted-by":"crossref","unstructured":"Calixto I, Liu Q, Campbell N (2017) Doubly-attentive decoder for multi-modal neural machine translation. In: Proceedings of the 55th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Vancouver, pp 1913\u20131924","DOI":"10.18653\/v1\/P17-1175"},{"key":"9250_CR33","doi-asserted-by":"crossref","unstructured":"Calixto I, Rios M, Aziz W (2019) Latent variable model for multi-modal translation. In: Proceedings of the 57th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Florence, pp 6392\u20136405","DOI":"10.18653\/v1\/P19-1642"},{"key":"9250_CR34","unstructured":"Carreira J, Noland E, Banki-Horvath A, Hillier C, Zisserman A (2018) A short note about Kinetics-600. Computing research repository arXiv:1808.01340"},{"issue":"1","key":"9250_CR35","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1023\/A:1007379606734","volume":"28","author":"R Caruana","year":"1997","unstructured":"Caruana R (1997) Multitask learning. Mach Learn 28(1):41\u201375","journal-title":"Mach Learn"},{"key":"9250_CR36","doi-asserted-by":"crossref","unstructured":"Castilho S, Doherty S, Gaspari F, Moorkens J (2018) Approaches to human and machine translation quality assessment. In: Translation quality assessment: from principles to practice, machine translation: technologies and applications, Springer, Berlin, pp 9\u201338","DOI":"10.1007\/978-3-319-91241-7_2"},{"key":"9250_CR37","unstructured":"Cettolo M, Girardi C, Federico M (2012) WIT3: web inventory of transcribed and translated talks. In: Proceedings of the 16th conference of the European Association for Machine Translation (EAMT), European Association for Machine Translation (EAMT), Trento, pp 261\u2013268"},{"key":"9250_CR38","unstructured":"Cettolo M, Niehues J, St\u00fcker S, Bentivogli L, Cattoni R, Federico M (2016) The IWSLT 2016 evaluation campaign. In: Proceedings of the (2016) International workshop on spoken language translation (IWSLT), Tokyo"},{"key":"9250_CR39","unstructured":"Cettolo M, Federico M, Bentivogli L, Niehues J, St\u00fcker S, Sudoh K, Yoshino K, Federmann C (2017) Overview of the IWSLT 2017 evaluation campaign. In: Proceedings of the 2017 international workshop on spoken language translation (IWSLT), Tokyo, pp 2\u201314"},{"key":"9250_CR40","unstructured":"Chen X, Fang H, Lin TY, Vedantam R, Gupta S, Dollar P, Zitnick CL (2015) Microsoft COCO captions: data collection and evaluation server. Computing research repository arXiv:1504.00325"},{"key":"9250_CR41","doi-asserted-by":"crossref","unstructured":"Chen Y, Liu Y, Li V (2018) Zero-resource neural machine translation with multi-agent communication game. In: 32nd AAAI conference on artificial intelligence, association for the advancement of artificial intelligence (AAAI)","DOI":"10.1609\/aaai.v32i1.11976"},{"key":"9250_CR42","doi-asserted-by":"crossref","unstructured":"Cheng Y, Tu Z, Meng F, Zhai J, Liu Y (2018) Towards robust neural machine translation. In: Proceedings of the 56th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Melbourne, pp 1756\u20131766","DOI":"10.18653\/v1\/P18-1163"},{"key":"9250_CR43","unstructured":"Chesterman A, Wagner E (2002) Can theory help translators?. Routledge, a dialogue between the Ivory Tower and the Wordface"},{"key":"9250_CR44","doi-asserted-by":"crossref","unstructured":"Chiu C, Sainath TN, Wu Y, Prabhavalkar R, Nguyen P, Chen Z, Kannan A, Weiss RJ, Rao K, Gonina E, Jaitly N, Li B, Chorowski J, Bacchiani M (2018) State-of-the-art speech recognition with sequence-to-sequence models. 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP). Calgary, pp 4774\u20134778","DOI":"10.1109\/ICASSP.2018.8462105"},{"key":"9250_CR45","doi-asserted-by":"crossref","unstructured":"Cho K, van Merri\u00ebnboer B, Bahdanau D, Bengio Y (2014) On the properties of neural machine translation: encoder\u2013decoder approaches. In: Proceedings of SSST-8, 8th workshop on syntax, semantics and structure in statistical translation, Association for Computational Linguistics (ACL), Doha, pp 103\u2013111","DOI":"10.3115\/v1\/W14-4012"},{"key":"9250_CR46","unstructured":"Chung J, Gulcehre C, Cho K, Bengio Y (2014) Empirical evaluation of gated recurrent neural networks on sequence modeling. In: Proceedings of the 27th neural information processing systems conference (NeurIPS) workshop on deep learning, Montreal"},{"key":"9250_CR47","doi-asserted-by":"crossref","unstructured":"Chung JS, Senior A, Vinyals O, Zisserman A (2017) Lip reading sentences in the wild. 2017 IEEE conference on computer vision and pattern recognition (CVPR). Honolulu, Hawaii, pp 3444\u20133453","DOI":"10.1109\/CVPR.2017.367"},{"key":"9250_CR48","unstructured":"Clough P, Grubinger M, Deselaers T, Hanbury A, M\u00fcller H (2006) Overview of the ImageCLEF 2006 photographic retrieval and object annotation tasks. In: Proceedings of the 7th international conference on cross-language evaluation forum (CLEF), Springer, Alicante, pp 579\u2013594"},{"key":"9250_CR49","volume-title":"Modulating and attending the source image during encoding improves multimodal translation. In: NIPS Workshop on Visually Grounded Interaction and Language (ViGIL)","author":"JB Delbrouck","year":"2017","unstructured":"Delbrouck JB, Dupont S (2017a) Modulating and attending the source image during encoding improves multimodal translation. In: NIPS Workshop on Visually Grounded Interaction and Language (ViGIL). Long Beach, California"},{"key":"9250_CR50","unstructured":"Delbrouck JB, Dupont S (2017b) Multimodal compact bilinear pooling for multimodal neural machine translation. Computing research repository arXiv:1703.08084"},{"key":"9250_CR51","unstructured":"Delbrouck JB, Dupont S (2018) UMONS submission for WMT18 multimodal translation task. In: Proceedings of the 3rd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Belgium, pp 643\u2013647"},{"key":"9250_CR52","unstructured":"Delbrouck JB, Dupont S (2019) Adversarial reconstruction for multi-modal machine translation. Computing research repository arXiv:1910.02766"},{"key":"9250_CR53","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009) Imagenet: a large-scale hierarchical image database. In: Proceedings of the IEEE conference on computer vision and pattern recognition, IEEE, pp 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"9250_CR54","doi-asserted-by":"crossref","unstructured":"Denkowski M, Lavie A (2014) Meteor universal: language specific translation evaluation for any target language. In: Proceedings of the 9th workshop on statistical machine translation, Association for Computational Linguistics (ACL), Baltimore, pp 376\u2013380","DOI":"10.3115\/v1\/W14-3348"},{"key":"9250_CR55","unstructured":"Di Gangi M, Negri M, Nguyen VN, Tebbifakhr A, Turchi M (2019a) Data augmentation for end-to-end speech translation: FBK\u00a0@\u00a0IWSLT\u201919. In: Proceedings of the 16th international workshop on spoken language translation (IWSLT), Hong Kong"},{"key":"9250_CR56","unstructured":"Di\u00a0Gangi MA, Cattoni R, Bentivogli L, Negri M, Turchi M (2019b) MuST-C: a multilingual speech translation corpus. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies (NAACL-HLT), Association for Computational Linguistics (ACL), Minneapolis, pp 2012\u20132017"},{"key":"9250_CR57","doi-asserted-by":"crossref","unstructured":"Di\u00a0Gangi MA, Negri M, Turchi M (2019c) Adapting transformer to end-to-end spoken language translation. In: Proceedings of interspeech, international speech communication association (ISCA), Graz, pp 1133\u20131137","DOI":"10.21437\/Interspeech.2019-3045"},{"key":"9250_CR58","doi-asserted-by":"crossref","unstructured":"Di Gangi MA, Negri M, Turchi M (2019d) One-to-many multilingual end-to-end speech translation. In: Proceedings of the (2019) IEEE workshop on automatic speech recognition and understanding (ASRU). Sentosa","DOI":"10.1109\/ASRU46091.2019.9004003"},{"key":"9250_CR59","unstructured":"Doherty S (2017) Issues in human and automatic translation quality assessment. In: Kenny D (ed) Human issues in translation technology: the IATIS yearbook. Routledge, pp 131\u2013148"},{"key":"9250_CR60","doi-asserted-by":"crossref","unstructured":"Dong D, Wu H, He W, Yu D, Wang H (2015) Multi-task learning for multiple language translation. In: Proceedings of the 53rd annual meeting of the association for computational linguistics (ACL) and the 7th international joint conference on natural language processing (IJCNLP), Association for Computational Linguistics (ACL), Beijing, pp 1723\u20131732","DOI":"10.3115\/v1\/P15-1166"},{"key":"9250_CR61","doi-asserted-by":"crossref","unstructured":"Dong L, Xu S, Xu B (2018) Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition. 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, Calgary, pp 5884\u20135888","DOI":"10.1109\/ICASSP.2018.8462506"},{"key":"9250_CR62","unstructured":"Drugan J (2013) Quality in professional translation: assessment and improvement. Continuum Advances in Translation, Bloomsbury Academic"},{"key":"9250_CR63","doi-asserted-by":"crossref","unstructured":"Duong L, Anastasopoulos A, Chiang D, Bird S, Cohn T (2016) An attentional model for speech translation without transcription. In: Proceedings of the 2016 conference of the North American chapter of the association for computational linguistics: human language technologies (NAACL-HLT), Association for Computational Linguistics (ACL), San Diego, pp 949\u2013959","DOI":"10.18653\/v1\/N16-1109"},{"key":"9250_CR64","doi-asserted-by":"crossref","unstructured":"Duselis J, Hutt M, Gwinnup J, Davis J, Sandvick J (2017) The AFRL-OSU WMT17 multimodal translation system: an image processing approach. In: Proceedings of the 2nd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Copenhagen, pp 445\u2013449","DOI":"10.18653\/v1\/W17-4748"},{"key":"9250_CR65","doi-asserted-by":"crossref","unstructured":"Dutta\u00a0Chowdhury K, Elliott D (2019) Understanding the effect of textual adversaries in multimodal machine translation. In: Proceedings of the Beyond Vision and LANguage: inTEgrating Real-world kNowledge (LANTERN), Association for Computational Linguistics (ACL), Hong Kong, pp 35\u201340","DOI":"10.18653\/v1\/D19-6406"},{"key":"9250_CR66","doi-asserted-by":"crossref","unstructured":"Elaraby M, Tawfik AY, Khaled M, Hassan H, Osama A (2018) Gender aware spoken language translation applied to English\u2013Arabic. In: Proceedings of the 2nd international conference on natural language and speech processing (ICNLSP), IEEE, Algiers, pp 1\u20136","DOI":"10.1109\/ICNLSP.2018.8374387"},{"key":"9250_CR67","doi-asserted-by":"crossref","unstructured":"Elliott D (2018) Adversarial evaluation of multimodal machine translation. In: Proceedings of the 2018 conference on empirical methods in natural language processing (EMNLP), Association for Computational Linguistics (ACL), pp 2974\u20132978","DOI":"10.18653\/v1\/D18-1329"},{"key":"9250_CR68","unstructured":"Elliott D, K\u00e1d\u00e1r \u00c1 (2017) Imagination improves multimodal translation. In: Proceedings of the 8th international joint conference on natural language processing (IJCNLP), Asian Federation of Natural Language Processing, Taipei, pp 130\u2013141"},{"key":"9250_CR69","unstructured":"Elliott D, Frank S, Hasler E (2015) Multi-language image description with neural sequence models. Computing research repository arXiv:1510.04709"},{"key":"9250_CR70","doi-asserted-by":"crossref","unstructured":"Elliott D, Frank S, Sima\u2019an K, Specia L (2016) Multi30k: multilingual English-German image descriptions. In: Proceedings of the 5th workshop on vision and language, Association for Computational Linguistics (ACL), Berlin, pp 70\u201374","DOI":"10.18653\/v1\/W16-3210"},{"key":"9250_CR71","doi-asserted-by":"crossref","unstructured":"Elliott D, Frank S, Barrault L, Bougares F, Specia L (2017) Findings of the second shared task on multimodal machine translation and multilingual image description. In: Proceedings of the 2nd conference on machine translation, Association for Computational Linguistics (ACL), Copenhagen, pp 215\u2013233","DOI":"10.18653\/v1\/W17-4718"},{"key":"9250_CR72","unstructured":"Federmann C, Lewis WD (2016) Microsoft speech language translation (MSLT) corpus: the IWSLT 2016 release for English, French and German. In: Proceedings of the 13th international workshop on spoken language translation (IWSLT), Seattle"},{"key":"9250_CR73","unstructured":"Federmann C, Lewis WD (2017) The Microsoft speech language translation (MSLT) corpus for Chinese and Japanese: conversational test data for machine translation and speech recognition. In: Proceedings of the machine translation summit XVI (MT Summit), Nagoya, pp 72\u201385"},{"key":"9250_CR74","unstructured":"Finn C, Abbeel P, Levine S (2017) Model-agnostic meta-learning for fast adaptation of deep networks. In: Precup D, Teh YW (eds) Proceedings of the 34th international conference on machine learning (ICML), PMLR, Sydney, proceedings of machine learning research, vol\u00a070, pp 1126\u20131135"},{"key":"9250_CR75","doi-asserted-by":"crossref","unstructured":"Firat O, Sankaran B, Al-onaizan Y, Yarman\u00a0Vural FT, Cho K (2016) Zero-resource translation with multi-lingual neural machine translation. In: Proceedings of the 2016 conference on empirical methods in natural language processing (EMNLP), Association for Computational Linguistics (ACL), Austin, pp 268\u2013277","DOI":"10.18653\/v1\/D16-1026"},{"key":"9250_CR76","doi-asserted-by":"crossref","unstructured":"Fomicheva M, Specia L (2016) Reference bias in monolingual machine translation evaluation. In: Proceedings of the 54th annual meeting of the Association for Computational Linguistics (ACL), Association for Computational Linguistics (ACL), Berlin, ACL, pp 77\u201382","DOI":"10.18653\/v1\/P16-2013"},{"issue":"03","key":"9250_CR77","doi-asserted-by":"publisher","first-page":"393","DOI":"10.1017\/S1351324918000074","volume":"24","author":"S Frank","year":"2018","unstructured":"Frank S, Elliott D, Specia L (2018) Assessing multilingual multimodal image description: studies of native speaker preferences and translator choices. Nat Lang Eng 24(03):393\u2013413","journal-title":"Nat Lang Eng"},{"key":"9250_CR78","doi-asserted-by":"crossref","unstructured":"Fukui A, Park DH, Yang D, Rohrbach A, Darrell T, Rohrbach M (2016) Multimodal compact bilinear pooling for visual question answering and visual grounding. In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics (ACL), Austin, pp 457\u2013468","DOI":"10.18653\/v1\/D16-1044"},{"key":"9250_CR79","unstructured":"Gehring J, Auli M, Grangier D, Yarats D, Dauphin YN (2017) Convolutional sequence to sequence learning. In: Proceedings of the 34th international conference on machine learning (ICML), JMLR.org, Sydney, ICML\u201917, pp 1243\u20131252"},{"key":"9250_CR80","doi-asserted-by":"crossref","unstructured":"Gella S, Sennrich R, Keller F, Lapata M (2017) Image pivoting for learning multilingual multimodal representations. In: Proceedings of the 2017 conference on empirical methods in natural language processing (EMNLP), Association for Computational Linguistics (ACL), Copenhagen, pp 2839\u20132845","DOI":"10.18653\/v1\/D17-1303"},{"key":"9250_CR81","doi-asserted-by":"crossref","unstructured":"Gella S, Elliott D, Keller F (2019) Cross-lingual visual verb sense disambiguation. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies (NAACL-HLT), Association for Computational Linguistics (ACL), Minneapolis, pp 1998\u20132004","DOI":"10.18653\/v1\/N19-1200"},{"key":"9250_CR82","doi-asserted-by":"crossref","unstructured":"Ghahremani P, BabaAli B, Povey D, Riedhammer K, Trmal J, Khudanpur S (2014) A pitch extraction algorithm tuned for automatic speech recognition. 2014 IEEE international conference on acoustics, speech and signal processing (ICASSP). Florence, pp 2494\u20132498","DOI":"10.1109\/ICASSP.2014.6854049"},{"key":"9250_CR83","doi-asserted-by":"crossref","unstructured":"Girshick R, Donahue J, Darrell T, Malik J (2014) Rich feature hierarchies for accurate object detection and semantic segmentation. In: The IEEE conference on computer vision and pattern recognition (CVPR), Columbus","DOI":"10.1109\/CVPR.2014.81"},{"key":"9250_CR84","unstructured":"Graham Y, Baldwin T, Moffat A, Zobel J (2013) Continuous measurement scales in human evaluation of machine translation. In: Proceedings of the 7th linguistic annotation workshop and interoperability with discourse, association for computational linguistics (ACL), Sofia, pp 33\u201341"},{"key":"9250_CR85","doi-asserted-by":"crossref","unstructured":"Graves A, Schmidhuber J (2005) Framewise phoneme classification with bidirectional LSTM networks. In: Proceedings. 2005 IEEE international joint conference on neural networks, 2005., IEEE, Montreal, vol\u00a04, pp 2047\u20132052","DOI":"10.1109\/IJCNN.2005.1556215"},{"key":"9250_CR86","doi-asserted-by":"crossref","unstructured":"Graves A, Ar Mohamed, Hinton G (2013) Speech recognition with deep recurrent neural networks. 2013 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, Vancouver, pp 6645\u20136649","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"9250_CR87","doi-asserted-by":"crossref","unstructured":"Gr\u00f6nroos SA, Huet B, Kurimo M, Laaksonen J, Merialdo B, Pham P, Sj\u00f6berg M, Sulubacak U, Tiedemann J, Troncy R, V\u00e1zquez R (2018) The MeMAD submission to the WMT18 multimodal translation task. In: Proceedings of the 3rd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Belgium, pp 609\u2013617","DOI":"10.18653\/v1\/W18-6439"},{"key":"9250_CR88","unstructured":"Grubinger M, Clough P, M\u00fcller H, Deselaers T (2006) The IAPR TC-12 benchmark: a new evaluation resource for visual information systems. In: Proceedings of the OntoImage workshop on language resources for content-based image retrieval, Genoa, pp 13\u201323"},{"key":"9250_CR89","unstructured":"Guzman F, Sajjad H, Vogel S, Abdelali A (2013) The AMARA corpus: building resources for translating the web\u2019s educational content. In: Proceedings of the 10th international workshop on spoken language translation (IWSLT), Heidelberg"},{"key":"9250_CR90","doi-asserted-by":"crossref","unstructured":"Gwinnup J, Sandvick J, Hutt M, Erdmann G, Duselis J, Davis J (2018) The AFRL-Ohio State WMT18 multimodal system: combining visual with traditional. In: Proceedings of the 3rd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Belgium, pp 618\u2013621","DOI":"10.18653\/v1\/W18-6440"},{"issue":"5","key":"9250_CR93","doi-asserted-by":"publisher","first-page":"1116","DOI":"10.1109\/JPROC.2012.2236631","volume":"101","author":"X He","year":"2013","unstructured":"He X, Deng L (2013) Speech-centric information processing: an optimization-oriented approach. Proc IEEE 101(5):1116\u20131135","journal-title":"Proc IEEE"},{"key":"9250_CR94","unstructured":"He X, Deng L, Acero A (2011) Why word error rate is not a good metric for speech recognizer training for the speech translation task? 2011 IEEE international conference on acoustics, speech and signal processing (ICASSP). Prague, pp 5632\u20135635"},{"key":"9250_CR91","doi-asserted-by":"crossref","unstructured":"He K, Xiangyu Z, Shaoqing R, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), Las Vegas, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"9250_CR92","doi-asserted-by":"crossref","unstructured":"He K, Gkioxari G, Doll\u00e1r P, Girshick R (2017) Mask R-CNN. In: Proceedings of the 2017 IEEE international conference on computer vision (ICCV), Venice, pp 2980\u20132988","DOI":"10.1109\/ICCV.2017.322"},{"key":"9250_CR95","doi-asserted-by":"crossref","unstructured":"Helcl J, Libovick\u00fd J (2017) CUNI system for the WMT17 multimodal translation task. In: Proceedings of the 2nd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Copenhagen, pp 450\u2013457","DOI":"10.18653\/v1\/W17-4749"},{"key":"9250_CR96","unstructured":"Helcl J, Libovick\u00fd J, Kocmi T, Musil T, C\u00edfka O, Vari\u0161 D, Bojar O (2018a) Neural Monkey: the current state and beyond. In: Proceedings of the 13th conference of the association for machine translation in the Americas (AMTA), Association for Machine Translation in the Americas, Boston, pp 168\u2013176"},{"key":"9250_CR97","doi-asserted-by":"crossref","unstructured":"Helcl J, Libovick\u00fd J, Varis D (2018b) CUNI system for the WMT18 multimodal translation task. In: Proceedings of the 3rd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Belgium, pp 622\u2013629","DOI":"10.18653\/v1\/W18-6441"},{"key":"9250_CR98","unstructured":"Hieber F, Domhan T, Denkowski M, Vilar D, Sokolov A, Clifton A, Post M (2017) Sockeye: a toolkit for neural machine translation. Computing Research Repository arXiv:1712.05690"},{"key":"9250_CR99","doi-asserted-by":"crossref","unstructured":"Hitschler J, Schamoni S, Riezler S (2016) Multimodal pivots for image caption translation. In: Proceedings of the 54th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Berlin, pp 2399\u20132409","DOI":"10.18653\/v1\/P16-1227"},{"issue":"8","key":"9250_CR100","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735\u20131780","journal-title":"Neural Comput"},{"key":"9250_CR101","doi-asserted-by":"crossref","unstructured":"Huang PY, Liu F, Shiang SR, Oh J, Dyer C (2016) Attention-based multimodal neural machine translation. In: Proceedings of the 1st conference on machine translation, Association for Computational Linguistics (ACL), Berlin, vol\u00a02, pp 639\u2013645","DOI":"10.18653\/v1\/W16-2360"},{"key":"9250_CR102","doi-asserted-by":"crossref","unstructured":"Inaguma H, Duh K, Kawahara T, Watanabe S (2019a) Multilingual end-to-end speech translation. In: Proceedings of the (2019) IEEE workshop on automatic speech recognition and understanding (ASRU). Sentosa, Singapore","DOI":"10.1109\/ASRU46091.2019.9003832"},{"key":"9250_CR103","unstructured":"Inaguma H, Kiyono S, Soplin NEY, Suzuki J, Duh K, Watanabe S (2019b) ESPnet How2 speech translation system for IWSLT 2019: Pre-training, knowledge distillation, and going deeper. In: Proceedings of the 16th international workshop on spoken language translation (IWSLT), Hong Kong"},{"key":"9250_CR104","doi-asserted-by":"crossref","unstructured":"Indurthi S, Han H, Lakumarapu NK, Lee B, Chung I, Kim S, Kim C (2020) End-end speech-to-text translation with modality agnostic meta-learning. In: 2020 ieee international conference on acoustics, speech and signal processing (ICASSP), Barcelona, pp 7904\u20137908","DOI":"10.1109\/ICASSP40776.2020.9054759"},{"key":"9250_CR105","unstructured":"Ioffe S, Szegedy C (2015) Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: Proceedings of the 32nd international conference on machine learning (ICML), Lille, pp 448\u2013456"},{"key":"9250_CR106","doi-asserted-by":"crossref","unstructured":"Ive J, Madhyastha P, Specia L (2019) Distilling translations with visual awareness. In: Proceedings of the 57th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Florence, pp 6525\u20136538","DOI":"10.18653\/v1\/P19-1653"},{"key":"9250_CR107","doi-asserted-by":"crossref","unstructured":"Jansen D, Alcala A, Guzman F (2014) AMARA: a sustainable, global solution for accessibility, powered by communities of volunteers. Universal access in human-computer interaction. Springer, Design for all and accessibility practice, pp 401\u2013411","DOI":"10.1007\/978-3-319-07509-9_38"},{"key":"9250_CR108","doi-asserted-by":"crossref","unstructured":"Jia Y, Johnson M, Macherey W, Weiss RJ, Cao Y, Chiu CC, Ari N, Laurenzo S, Wu Y (2019) Leveraging weakly supervised data to improve end-to-end speech-to-text translation. 2019 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, Brighton, pp 7180\u20137184","DOI":"10.1109\/ICASSP.2019.8683343"},{"key":"9250_CR109","doi-asserted-by":"publisher","first-page":"339","DOI":"10.1162\/tacl_a_00065","volume":"5","author":"M Johnson","year":"2017","unstructured":"Johnson M, Schuster M, Le QV, Krikun M, Wu Y, Chen Z, Thorat N, Vi\u00e9gas F, Wattenberg M, Corrado G, Hughes M, Dean J (2017) Google\u2019s multilingual neural machine translation system: enabling zero-shot translation. Trans Assoc Comput Linguistics 5:339\u2013351","journal-title":"Trans Assoc Comput Linguistics"},{"key":"9250_CR110","doi-asserted-by":"crossref","unstructured":"Junczys-Dowmunt M, Grundkiewicz R, Dwojak T, Hoang H, Heafield K, Neckermann T, Seide F, Germann U, Aji AF, Bogoychev N, Martins AFT, Birch A (2018) Marian: fast neural machine translation in C++. In: Proceedings of the 56th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Melbourne, pp 116\u2013121","DOI":"10.18653\/v1\/P18-4020"},{"key":"9250_CR111","doi-asserted-by":"crossref","unstructured":"K\u00e1d\u00e1r \u00c1, Elliott D, C\u00f4t\u00e9 MA, Chrupa\u0142a G, Alishahi A (2018) Lessons learned in multilingual grounded language learning. In: Proceedings of the 22nd conference on computational natural language learning (CoNLL), Association for Computational Linguistics (ACL), Brussels, pp 402\u2013412","DOI":"10.18653\/v1\/K18-1039"},{"key":"9250_CR112","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1016\/j.cviu.2017.06.005","volume":"163","author":"K Kafle","year":"2017","unstructured":"Kafle K, Kanan C (2017) Visual question answering: datasets, algorithms, and future challenges. Comput Vis Image Underst 163:3\u201320","journal-title":"Comput Vis Image Underst"},{"key":"9250_CR113","unstructured":"Kalchbrenner N, Blunsom P (2013) Recurrent continuous translation models. In: Proceedings of the 2013 conference on empirical methods in natural language processing (EMNLP), Association for Computational Linguistics (ACL), Seattle, pp 1700\u20131709"},{"key":"9250_CR114","unstructured":"Kay W, Carreira J, Simonyan K, Zhang B, Hillier C, Vijayanarasimhan S, Viola F, Green T, Back T, Natsev P, Suleyman M, Zisserman A (2017) The Kinetics human action video dataset. Computing Research Repository arXiv:1705.06950"},{"key":"9250_CR115","unstructured":"Kiros R, Salakhutdinov R, Zemel R (2014) Multimodal neural language models. In: Proceedings of the 31st international conference on machine learning (ICML), Beijing"},{"key":"9250_CR116","doi-asserted-by":"crossref","unstructured":"Klein G, Kim Y, Deng Y, Senellart J, Rush A (2017) OpenNMT: open-source toolkit for neural machine translation. In: Proceedings of the 55th annual meeting of the association for computational linguistics, Association for Computational Linguistics (ACL), Vancouver, pp 67\u201372","DOI":"10.18653\/v1\/P17-4012"},{"key":"9250_CR117","unstructured":"Kocabiyikoglu AC, Besacier L, Kraif O (2018) Augmenting Librispeech with French translations: a multimodal corpus for direct speech translation evaluation. In: Proceedings of the 11th conference on language resources and evaluation (LREC), European Language Resources Association (ELRA), Miyazaki"},{"key":"9250_CR118","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511815829","volume-title":"Statistical machine translation","author":"P Koehn","year":"2009","unstructured":"Koehn P (2009) Statistical machine translation. Cambridge University Press, Cambridge"},{"key":"9250_CR119","unstructured":"Koehn P, Zens R, Dyer C, Bojar O, Constantin A, Herbst E, Hoang H, Birch A, Callison-Burch C, Federico M, Bertoldi N, Cowan B, Shen W, Moran C (2007) Moses: open source toolkit for statistical machine translation. In: Proceedings of the 45th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Prague"},{"key":"9250_CR120","doi-asserted-by":"crossref","unstructured":"Kreutzer J, Bastings J, Riezler S (2019) Joey NMT: a minimalist NMT toolkit for novices. In: Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP\u2013IJCNLP): system demonstrations, Association for Computational Linguistics (ACL), Hong Kong, pp 109\u2013114","DOI":"10.18653\/v1\/D19-3019"},{"key":"9250_CR121","unstructured":"Lala C, Specia L (2018) Multimodal lexical translation. In: Proceedings of the 11th international conference on language resources and evaluation (LREC), Miyazaki"},{"key":"9250_CR122","doi-asserted-by":"crossref","unstructured":"Lala C, Madhyastha PS, Scarton C, Specia L (2018) Sheffield submissions for WMT18 multimodal translation shared task. In: Proceedings of the 3rd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Belgium, pp 630\u2013637","DOI":"10.18653\/v1\/W18-6442"},{"key":"9250_CR123","doi-asserted-by":"crossref","unstructured":"Lavie A, Agarwal A (2007) METEOR: an automatic metric for MT evaluation with high levels of correlation with human judgments. In: Proceedings of the 2nd workshop on statistical machine translation - StatMT \u201907, Association for Computational Linguistics (ACL), Prague, pp 228\u2013231","DOI":"10.3115\/1626355.1626389"},{"key":"9250_CR124","doi-asserted-by":"crossref","unstructured":"Lavie A, Waibel A, Levin L, Finke M, Gates D, Gavalda M, Zeppenfeld T, Zhan P (1997) JANUS-III: Speech-to-speech translation in multiple languages. In: 1997 IEEE international conference on acoustics, speech, and signal processing (ICASSP), IEEE Comput. Soc. Press, Munich, vol\u00a01, pp 99\u2013102","DOI":"10.1109\/ICASSP.1997.599557"},{"key":"9250_CR126","doi-asserted-by":"crossref","unstructured":"Li X, Lan W, Dong J, Liu H (2016) Adding Chinese captions to images. In: Proceedings of the 2016 ACM on international conference on multimedia retrieval-ICMR\u201916, ACM Press, New York, pp 271\u2013275","DOI":"10.1145\/2911996.2912049"},{"key":"9250_CR125","doi-asserted-by":"crossref","unstructured":"Li J, Lavrukhin V, Ginsburg B, Leary R, Kuchaiev O, Cohen JM, Nguyen H, Gadde RT (2019) Jasper: an end-to-end convolutional neural acoustic model. In: Proceedings of Interspeech, Graz, pp 71\u201375","DOI":"10.21437\/Interspeech.2019-1819"},{"key":"9250_CR127","unstructured":"Libovick\u00fd J (2019) Multimodality in machine translation. PhD thesis, Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics, Prague"},{"key":"9250_CR128","doi-asserted-by":"crossref","unstructured":"Libovick\u00fd J, Helcl J (2017) Attention strategies for multi-source sequence-to-sequence learning. In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL), Association for Computational Linguistics (ACL), Vancouver, pp 196\u2013202","DOI":"10.18653\/v1\/P17-2031"},{"key":"9250_CR129","doi-asserted-by":"crossref","unstructured":"Libovick\u00fd J, Helcl J, Tlust\u00fd M, Bojar O, Pecina P (2016) CUNI system for WMT16 automatic post-editing and multimodal translation tasks. In: Proceedings of the 1st Conference on Machine Translation (WMT), Association for Computational Linguistics (ACL), Berlin, pp 646\u2013654","DOI":"10.18653\/v1\/W16-2361"},{"key":"9250_CR130","doi-asserted-by":"crossref","unstructured":"Libovick\u00fd J, Helcl J, Mare\u010dek D (2018) Input combination strategies for multi-source transformer decoder. In: Proceedings of the 3rd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Belgium, pp 253\u2013260","DOI":"10.18653\/v1\/W18-6326"},{"key":"9250_CR131","unstructured":"Lin M, Chen Q, Yan S (2014a) Network in network. In: Proceedings of the 2nd international conference on learning representations (ICLR), Scottsdale"},{"key":"9250_CR132","doi-asserted-by":"crossref","unstructured":"Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014b) Microsoft COCO: Common Objects in Context. In: Fleet D, Pajdla T, Schiele B, Tuytelaars T (eds) Proceedings of the 13th European Conference on Computer Vision (ECCV), Springer International Publishing, Zurich, vol 8693, pp 740\u2013755","DOI":"10.1007\/978-3-319-10602-1_48"},{"issue":"32","key":"9250_CR133","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1109\/MSP.2014.2359987","volume":"3","author":"ZH Ling","year":"2015","unstructured":"Ling ZH, Kang SY, Zen H, Senior A, Schuster M, Qian XJ, Meng HM, Deng L (2015) Deep learning for acoustic modeling in parametric speech generation: a systematic review of existing techniques and future trends. IEEE Signal Process Mag 3(32):35\u201352","journal-title":"IEEE Signal Process Mag"},{"key":"9250_CR134","unstructured":"Lison P, Tiedemann J (2016) OpenSubtitles2016: extracting large parallel corpora from movie and TV subtitles. In: Chair) NCC, Choukri K, Declerck T, Goggi S, Grobelnik M, Maegaard B, Mariani J, Mazo H, Moreno A, Odijk J, Piperidis S (eds) Proceedings of the 10th international conference on language resources and evaluation (LREC), European Language Resources Association (ELRA), Portoro\u017e"},{"key":"9250_CR135","unstructured":"Liu D, Liu J, Guo W, Xiong S, Ma Z, Song R, Wu C, Liu Q (2018) The USTC-NEL speech translation system at IWSLT 2018. In: Proceedings of the 15th international workshop on spoken language translation (IWSLT), pp 70\u201375"},{"key":"9250_CR136","doi-asserted-by":"crossref","unstructured":"Liu Y, Xiong H, He Z, Zhang J, Wu H, Wang H, Zong C (2019) End-to-end speech translation with knowledge distillation. In: Proceedings of Interspeech, Graz","DOI":"10.21437\/Interspeech.2019-2582"},{"key":"9250_CR137","unstructured":"Luong T, Le QV, Sutskever I, Vinyals O, Kaiser L (2016) Multi-task sequence to sequence learning. In: Proceedings of the 4th international conference on learning representations (ICLR), San Juan"},{"key":"9250_CR138","doi-asserted-by":"crossref","unstructured":"L\u00fcscher C, Beck E, Irie K, Kitza M, Michel W, Zeyer A, Schl\u00fcter R, Ney H (2019) RWTH ASR systems for LibriSpeech: hybrid vs attention. In: Proceedings of interspeech, Graz, pp 231\u2013235","DOI":"10.21437\/Interspeech.2019-1780"},{"key":"9250_CR139","doi-asserted-by":"crossref","unstructured":"Ma M, Li D, Zhao K, Huang L (2017) OSU multimodal machine translation system report. In: Proceedings of the 2nd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Copenhagen, pp 465\u2013469","DOI":"10.18653\/v1\/W17-4751"},{"key":"9250_CR140","doi-asserted-by":"crossref","unstructured":"Ma Q, Bojar O, Graham Y (2018) Results of the WMT18 metrics shared task: Both characters and embeddings achieve good performance. In: Proceedings of the 3rd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Belgium, pp 682\u2013701","DOI":"10.18653\/v1\/W18-6450"},{"key":"9250_CR141","doi-asserted-by":"crossref","unstructured":"Ma Q, Wei J, Bojar O, Graham Y (2019) Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges. In: Proceedings of the 4th conference on machine translation (WMT), Association for Computational Linguistics (ACL), Florence, pp 62\u201390","DOI":"10.18653\/v1\/W19-5302"},{"key":"9250_CR143","doi-asserted-by":"crossref","unstructured":"Madhyastha PS, Wang J, Specia L (2017) Sheffield MultiMT: using object posterior predictions for multimodal machine translation. In: Proceedings of the 2nd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Copenhagen, pp 470\u2013476","DOI":"10.18653\/v1\/W17-4752"},{"key":"9250_CR142","doi-asserted-by":"crossref","unstructured":"Madhyastha P, Wang J, Specia L (2019) VIFIDEL: evaluating the visual fidelity of image descriptions. In: Proceedings of the 57th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Florence, pp 6539\u20136550","DOI":"10.18653\/v1\/P19-1654"},{"key":"9250_CR144","unstructured":"Mao J, Xu W, Yang Y, Wang J, Huang Z, Yuille A (2015) Deep captioning with multimodal recurrent neural networks (m-rnn). In: Proceedings of the 3rd international conference on learning representations (ICLR), Banff"},{"key":"9250_CR145","doi-asserted-by":"crossref","unstructured":"Matusov E, Kanthak S, Ney H (2006) Integrating speech recognition and machine translation: Where do we stand? In: 2006 IEEE international conference on acoustics speech and signal processing proceedings, Toulouse, vol\u00a05, pp 1217\u20131220","DOI":"10.1109\/ICASSP.2006.1661501"},{"key":"9250_CR146","doi-asserted-by":"crossref","unstructured":"Mikolov T, Karafi\u00e1t M, Burget L, Cernock\u00fd J, Khudanpur S (2010) Recurrent neural network based language model. In: Kobayashi T, Hirose K, Nakamura S (eds) Proceedings of interspeech, ISCA, Makuhari, Chiba, pp 1045\u20131048","DOI":"10.21437\/Interspeech.2010-343"},{"key":"9250_CR147","doi-asserted-by":"crossref","unstructured":"Miyazaki T, Shimizu N (2016) Cross-lingual image caption generation. In: Proceedings of the 54th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Berlin, pp 1780\u20131790","DOI":"10.18653\/v1\/P16-1168"},{"key":"9250_CR148","unstructured":"Mogadala A, Kalimuthu M, Klakow D (2019) Trends in integration of vision and language research: A survey of tasks, datasets, and methods. Computing research repository arXiv:1907.09358"},{"key":"9250_CR149","doi-asserted-by":"crossref","unstructured":"Mohamed A, Hinton G, Penn G (2012) Understanding how deep belief networks perform acoustic modelling. 2012 IEEE international conference on acoustics, speech and signal processing (ICASSP). Kyoto, pp 4273\u20134276","DOI":"10.1109\/ICASSP.2012.6288863"},{"key":"9250_CR150","unstructured":"Morimoto T (1990) Automatic interpreting telephony research at ATR. In: Proceedings of a workshop on machine translation, UMIST, Manchester"},{"issue":"1\u20132","key":"9250_CR151","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1007\/s10590-017-9197-z","volume":"31","author":"H Nakayama","year":"2017","unstructured":"Nakayama H, Nishida N (2017) Zero-resource machine translation by multimodal encoder\u2013decoder network with multimedia pivot. Mach Transl 31(1\u20132):49\u201364","journal-title":"Mach Transl"},{"key":"9250_CR152","doi-asserted-by":"crossref","unstructured":"Ney H (1999) Speech translation: coupling of recognition and translation. In: 1999 IEEE international conference on acoustics, speech, and signal processing (ICASSP), IEEE, Phoenix, Arizona, vol\u00a01, pp 517\u2013520","DOI":"10.1109\/ICASSP.1999.758176"},{"key":"9250_CR153","unstructured":"Niehues J, Cattoni R, St\u00fcker S, Cettolo M, Turchi M, Federico M (2018) The IWSLT 2018 evaluation campaign. In: Proceedings of the (2018) International workshop on spoken language translation (IWSLT). Bruges"},{"key":"9250_CR154","unstructured":"Niehues J, Cattoni R, St\u00fcker S, Negri M, Turchi M, Ha TL, Salesky E, Sanabria R, Barrault L, Specia L, Federico M (2019) The IWSLT 2019 evaluation campaign. In: Proceedings of the 16th international workshop on spoken language translation (IWSLT)"},{"key":"9250_CR155","doi-asserted-by":"crossref","unstructured":"Och FJ (2003) Minimum error rate training in statistical machine translation. In: Proceedings of the 41st annual meeting on association for computational linguistics-volume 1, Association for Computational Linguistics (ACL), Stroudsburg, ACL\u201903, pp 160\u2013167","DOI":"10.3115\/1075096.1075117"},{"key":"9250_CR156","unstructured":"Osamura K, Kano T, Sakti S, Sudoh K, Nakamura S (2018) Using spoken word posterior features in neural machine translation. In: Proceedings of the 15th international workshop on spoken language translation (IWSLT), Bruges, pp 189\u2013195"},{"key":"9250_CR157","doi-asserted-by":"crossref","unstructured":"Ott M, Edunov S, Baevski A, Fan A, Gross S, Ng N, Grangier D, Auli M (2019) fairseq: A fast, extensible toolkit for sequence modeling. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics (NAACL-HLT), Association for Computational Linguistics (ACL), Minneapolis, pp 48\u201353","DOI":"10.18653\/v1\/N19-4009"},{"key":"9250_CR158","doi-asserted-by":"crossref","unstructured":"Panayotov V, Chen G, Povey D, Khudanpur S (2015) Librispeech: an ASR corpus based on public domain audio books. 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, South Brisbane, pp 5206\u20135210","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"9250_CR159","doi-asserted-by":"crossref","unstructured":"Papineni K, Roukos S, Ward T, Zhu WJ (2001) BLEU: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting on association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Philadelphia, pp 311\u2013318","DOI":"10.3115\/1073083.1073135"},{"key":"9250_CR160","unstructured":"Paul M, Federico M, St\u00fcker S (2010) Overview of the IWSLT 2010 evaluation campaign. In: Proceedings of the (2010) international workshop on spoken language translation (IWSLT). France"},{"key":"9250_CR161","unstructured":"Peitz S, Wiesler S, Nussbaum-Thom M, Ney H (2012) Spoken language translation using automatically transcribed text in training. In: Proceedings of the 9th international workshop on spoken language translation (IWSLT), Hong Kong, pp 276\u2013283"},{"key":"9250_CR162","doi-asserted-by":"crossref","unstructured":"Pham NQ, Nguyen TS, Ha TL, Hussain J, Schneider F, Niehues J, St\u00fcker S, Waibel A (2019) The iwslt 2019 kit speech translation system. In: Proceedings of the 16th international workshop on spoken language translation (IWSLT), Hong Kong","DOI":"10.18653\/v1\/2020.iwslt-1.4"},{"key":"9250_CR163","unstructured":"Pino J, Puzon L, Gu J, Ma X, McCarthy AD, Gopinath D (2019) Harnessing indirect training data for end-to-end automatic speech translation: tricks of the trade. In: Proceedings of the 16th international workshop on spoken language translation (IWSLT)"},{"key":"9250_CR164","unstructured":"Post M, Kumar G, Lopez A, Karakos D, Callison-Burch C, Khudanpur S (2013) Improved speech-to-text translation with the Fisher and Callhome Spanish-English speech translation corpus. In: Proceedings of the 10th international workshop on spoken language translation (IWSLT), Heidelberg"},{"key":"9250_CR165","doi-asserted-by":"crossref","DOI":"10.1002\/9781119825449","volume-title":"Communication acoustics: an introduction to speech, audio and psychoacoustics","author":"V Pulkki","year":"2015","unstructured":"Pulkki V, Karjalainen M (2015) Communication acoustics: an introduction to speech, audio and psychoacoustics. Wiley, Chichester"},{"key":"9250_CR166","doi-asserted-by":"crossref","unstructured":"Ramanathan V, Joulin A, Liang P, Fei-Fei L (2014) Linking people in videos with \u201ctheir\u201d names using coreference resolution. In: Proceedings of the 13th European conference on computer vision (ECCV), Springer, pp 95\u2013110","DOI":"10.1007\/978-3-319-10590-1_7"},{"key":"9250_CR167","doi-asserted-by":"crossref","unstructured":"Ramirez J, Gorriz JM, Segura JC (2007) Voice activity detection. Fundamentals and speech recognition system robustness. In: Grimm M, Kroschel K (eds) Robust speech, IntechOpen, Rijeka, chap\u00a01","DOI":"10.5772\/4740"},{"key":"9250_CR168","unstructured":"Rashtchian C, Young P, Hodosh M, Hockenmaier J (2010) Collecting image annotations using Amazon\u2019s Mechanical Turk. In: Proceedings of the workshop on creating speech and language data with Amazon\u2019s Mechanical Turk, Association for Computational Linguistics (ACL), pp 139\u2013147"},{"key":"9250_CR169","unstructured":"Ruiz N, Federico M (2014) Assessing the impact of speech recognition errors on machine translation quality. In: Proceedings of the 11th conference of the association for machine translation in the Americas (AMTA), Vancouver, pp 261\u2013274"},{"key":"9250_CR170","doi-asserted-by":"crossref","unstructured":"Ruiz N, Federico M (2015) Phonetically-oriented word error alignment for speech recognition error analysis in speech translation. In: Proceedings of the 2015 IEEE workshop on automatic speech recognition and understanding (ASRU), Scottsdale, Arizona, pp 296\u2013302","DOI":"10.1109\/ASRU.2015.7404808"},{"key":"9250_CR171","doi-asserted-by":"crossref","unstructured":"Ruiz N, Gangi MAD, Bertoldi N, Federico M (2017) Assessing the tolerance of neural machine translation systems against speech recognition errors. In: Proceedings of Interspeech, Stockholm, pp 2635\u20132639","DOI":"10.21437\/Interspeech.2017-1690"},{"issue":"3","key":"9250_CR172","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S, Huang Z, Karpathy A, Khosla A, Bernstein M, Berg AC, Fei-Fei L (2015) ImageNet large scale visual recognition challenge. IJCV 115(3):211\u2013252","journal-title":"IJCV"},{"key":"9250_CR173","doi-asserted-by":"crossref","unstructured":"Sainath TN, Weiss RJ, Senior A, Wilson KW, Vinyals O (2015) Learning the speech front-end with raw waveform cldnns. In: 16th annual conference of the international speech communication association (ISCA), Dresden","DOI":"10.21437\/Interspeech.2015-1"},{"key":"9250_CR174","doi-asserted-by":"crossref","unstructured":"Salesky E, Sperber M, Waibel A (2019) Fluent translations from disfluent speech in end-to-end speech translation. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies (NAACL-HLT), Association for Computational Linguistics (ACL), Minneapolis, pp 2786\u20132792","DOI":"10.18653\/v1\/N19-1285"},{"key":"9250_CR175","unstructured":"Sanabria R, Caglayan O, Palaskar S, Elliott D, Barrault L, Specia L, Metze F (2018) How2: a large-scale dataset for multimodal language understanding. In: NeurIPS, workshop on visually grounded interaction and language (ViGIL). Montreal"},{"key":"9250_CR176","doi-asserted-by":"crossref","unstructured":"Saon G, Soltau H, Nahamoo D, Picheny M (2013) Speaker adaptation of neural network acoustic models using i-vectors. In: Proceedings of the 2013 IEEE workshop on automatic speech recognition and understanding (ASRU), Olomouc, pp 55\u201359","DOI":"10.1109\/ASRU.2013.6707705"},{"key":"9250_CR177","unstructured":"Schneider F, Waibel A (2019) KIT\u2019s submission to the IWSLT 2019 shared task on text translation. In: Proceedings of the 16th international workshop on spoken language translation (IWSLT), Hong Kong"},{"issue":"11","key":"9250_CR178","doi-asserted-by":"publisher","first-page":"2673","DOI":"10.1109\/78.650093","volume":"45","author":"M Schuster","year":"1997","unstructured":"Schuster M, Paliwal KK (1997) Bidirectional recurrent neural networks. IEEE Trans Signal Process 45(11):2673\u20132681","journal-title":"IEEE Trans Signal Process"},{"key":"9250_CR179","doi-asserted-by":"crossref","unstructured":"Schwenk H, Dechelotte D, Gauvain JL (2006) Continuous space language models for statistical machine translation. In: Proceedings of the 2006 joint conference on computational linguistics (COLING) and annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Sydney, pp 723\u2013730","DOI":"10.3115\/1273073.1273166"},{"key":"9250_CR180","doi-asserted-by":"crossref","unstructured":"Sennrich R, Haddow B, Birch A (2016a) Improving neural machine translation models with monolingual data. In: Proceedings of the 54th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), pp 86\u201396","DOI":"10.18653\/v1\/P16-1009"},{"key":"9250_CR181","doi-asserted-by":"crossref","unstructured":"Sennrich R, Haddow B, Birch A (2016b) Neural machine translation of rare words with subword units. In: Proceedings of the 54th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Berlin, pp 1715\u20131725","DOI":"10.18653\/v1\/P16-1162"},{"key":"9250_CR182","doi-asserted-by":"crossref","unstructured":"Sennrich R, Firat O, Cho K, Birch-Mayne A, Haddow B, Hitschler J, Junczys-Dowmunt M, L\u00e4ubli S, Miceli Barone A, Mokry J, Nadejde M (2017) Nematus: a toolkit for neural machine translation. In: Proceedings of the conference of the European chapter of the association for computational linguistics (EACL): software demonstrations, Association for Computational Linguistics (ACL), Valencia, pp 65\u201368","DOI":"10.18653\/v1\/E17-3017"},{"key":"9250_CR183","doi-asserted-by":"crossref","unstructured":"Shah K, Wang J, Specia L (2016) Shef-multimodal: Grounding machine translation on images. In: Proceedings of the 1st conference on machine translation (WMT), Association for Computational Linguistics (ACL), Berlin, pp 660\u2013665","DOI":"10.18653\/v1\/W16-2363"},{"key":"9250_CR184","unstructured":"Shen J, Nguyen P, Wu Y, Chen Z et al (2019) Lingvo: a modular and scalable framework for sequence-to-sequence modeling. Computing research repository arXiv:1902.08295"},{"key":"9250_CR185","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. In: Proceedings of the 3rd international conference on learning representations (ICLR), Banff"},{"key":"9250_CR186","unstructured":"Snover M, Dorr B, Schwartz R, Micciulla L, Makhoul J (2006) A study of translation edit rate with targeted human annotation. In: Proceedings of the 7th conference of the association for machine translation in the Americas (AMTA), Cambridge, pp 223\u2013231"},{"key":"9250_CR187","doi-asserted-by":"crossref","unstructured":"Specia L, Frank S, Sima\u2019an K, Elliott D (2016) A shared task on multimodal machine translation and crosslingual image description. In: Proceedings of the 1st conference on machine translation: volume 2, shared task papers, Association for Computational Linguistics (ACL), Berlin, pp 543\u2013553","DOI":"10.18653\/v1\/W16-2346"},{"key":"9250_CR188","first-page":"55","volume-title":"Machine Translation Summit XVI","author":"L Specia","year":"2017","unstructured":"Specia L, Harris K, Blain F, Burchardt A, Macketanz V, Skadina I, Negri M, Turchi M (2017) Translation quality and productivity: a study on rich morphology languages. Machine Translation Summit XVI. Nagoya, Japan, pp 55\u201371"},{"key":"9250_CR189","doi-asserted-by":"crossref","unstructured":"Specia L, Blain F, Logacheva V, Astudillo RF, Martins A (2018) Findings of the WMT 2018 shared task on quality estimation. In: Proceedings of the 3rd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Belgium, pp 702\u2013722","DOI":"10.18653\/v1\/W18-6451"},{"key":"9250_CR190","doi-asserted-by":"crossref","unstructured":"Sperber M, Neubig G, Niehues J, Waibel A (2017a) Neural lattice-to-sequence models for uncertain inputs. In: Proceedings of the 2017 conference on empirical methods in natural language processing (EMNLP), Association for Computational Linguistics (ACL), Copenhagen, pp 1380\u20131389","DOI":"10.18653\/v1\/D17-1145"},{"key":"9250_CR191","unstructured":"Sperber M, Niehues J, Waibel A (2017b) Toward robust neural machine translation for noisy input sequences. In: Proceedings of the 14th international workshop on spoken language translation (IWSLT), Tokyo, pp 90\u201396"},{"key":"9250_CR192","doi-asserted-by":"publisher","first-page":"313","DOI":"10.1162\/tacl_a_00270","volume":"7","author":"M Sperber","year":"2019","unstructured":"Sperber M, Neubig G, Niehues J, Waibel A (2019) Attention-passing models for robust and data-efficient end-to-end speech translation. Trans Assoc Comput Linguistics 7:313\u2013325","journal-title":"Trans Assoc Comput Linguistics"},{"issue":"1","key":"9250_CR193","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1016\/j.heares.2009.03.012","volume":"258","author":"BE Stein","year":"2009","unstructured":"Stein BE, Stanford TR, Rowland BA (2009) The neural basis of multisensory integration in the midbrain: its organization and maturation. Hear Res 258(1):4\u201315","journal-title":"Hear Res"},{"key":"9250_CR194","unstructured":"Stoian MC, Bansal S, Goldwater S (2019) Analyzing ASR pretraining for low-resource speech-to-text translation. Computing research repository arXiv:1910.10762"},{"key":"9250_CR195","unstructured":"Sutskever I, Vinyals O, Le QV (2014) Sequence to sequence learning with neural networks. In: Proceedings of the 27th international conference on neural information processing systems (NeurIPS), MIT Press, Montreal, NIPS\u201914, pp 3104\u20133112"},{"key":"9250_CR196","doi-asserted-by":"crossref","unstructured":"Takezawa T, Morimoto T, Sagisaka Y, Campbell N, Iida H, Sugaya F, Yokoo A, Yamamoto S (1998) A Japanese-to-English speech translation system: ATR-MATRIX. In: Proceedings of the 5th international conference on spoken language processing (ICSLP), Sydney","DOI":"10.21437\/ICSLP.1998-581"},{"key":"9250_CR197","unstructured":"Tiedemann J (2012) Parallel data, tools and interfaces in OPUS. In: Calzolari N, Choukri K, Declerck T, Do\u011fan MU, Maegaard B, Mariani J, Moreno A, Odijk J, Piperidis S (eds) Proceedings of the 8th international conference on language resources and evaluation (LREC), European language resources association (ELRA), Istanbul"},{"key":"9250_CR198","unstructured":"Toyama J, Misono M, Suzuki M, Nakayama K, Matsuo Y (2016) Neural machine translation with latent semantic of image and text. Computing research repository arXiv:1611.08459"},{"key":"9250_CR199","doi-asserted-by":"crossref","unstructured":"Tsvetkov Y, Metze F, Dyer C (2014) Augmenting translation models with simulated acoustic confusions for improved spoken language translation. In: Proceedings of the 14th conference of the European chapter of the association for computational linguistics (EACL), Association for Computational Linguistics (ACL), Gothenburg, pp 616\u2013625","DOI":"10.3115\/v1\/E14-1065"},{"key":"9250_CR200","doi-asserted-by":"crossref","unstructured":"Unal ME, Citamak B, Yagcioglu S, Erdem A, Erdem E, Cinbis NI, Cakici R (2016) Tasviret: a benchmark dataset for automatic Turkish description generation from images. In: 2016 24th signal processing and communication application conference (SIU), pp 1977\u20131980","DOI":"10.1109\/SIU.2016.7496155"},{"key":"9250_CR201","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. In: Guyon I, Luxburg UV, Bengio S, Wallach H, Fergus R, Vishwanathan S, Garnett R (eds) Advances in neural information processing systems 30, Curran Associates, Inc., pp 5998\u20136008"},{"key":"9250_CR202","unstructured":"Vaswani A, Bengio S, Brevdo E, Chollet F, Gomez A, Gouws S, Jones L, Kaiser \u0141, Kalchbrenner N, Parmar N, Sepassi R, Shazeer N, Uszkoreit J (2018) Tensor2Tensor for neural machine translation. In: Proceedings of the 13th conference of the association for machine translation in the Americas (AMTA), Association for Machine Translation in the Americas, Boston, pp 193\u2013199"},{"key":"9250_CR203","doi-asserted-by":"crossref","unstructured":"Vidal E (1997) Finite-state speech-to-speech translation. In: 1997 IEEE international conference on acoustics, speech, and signal processing (ICASSP), IEEE, Munich, vol\u00a01, pp 111\u2013114","DOI":"10.1109\/ICASSP.1997.599563"},{"key":"9250_CR204","doi-asserted-by":"crossref","unstructured":"Vinyals O, Toshev A, Bengio S, Erhan D (2015) Show and tell: a neural image caption generator. In: Proceedings of the IEEE conference on computer vision and pattern recognition, IEEE, pp 3156\u20133164","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"9250_CR205","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/978-3-662-04230-4_1","volume-title":"Verbmobil: foundations of speech-to-speech translation","author":"W Wahlster","year":"2000","unstructured":"Wahlster W (2000) Mobile speech-to-speech translation of spontaneous dialogs: an overview of the final Verbmobil system. In: Wahlster W (ed) Verbmobil: foundations of speech-to-speech translation. Springer, Heidelberg, pp 3\u201321"},{"key":"9250_CR206","doi-asserted-by":"crossref","unstructured":"Wang A, Singh A, Michael J, Hill F, Levy O, Bowman S (2018a) GLUE: A multi-task benchmark and analysis platform for natural language understanding. In: Proceedings of the 2018 EMNLP workshop BlackboxNLP: analyzing and interpreting neural networks for NLP, association for computational linguistics (ACL), Brussels, pp 353\u2013355","DOI":"10.18653\/v1\/W18-5446"},{"key":"9250_CR208","doi-asserted-by":"crossref","unstructured":"Wang X, Pham H, Dai Z, Neubig G (2018b) SwitchOut: an efficient data augmentation algorithm for neural machine translation. In: Proceedings of the 2018 conference on empirical methods in natural language processing (EMNLP), Association for Computational Linguistics (ACL), Brussels, pp 856\u2013861","DOI":"10.18653\/v1\/D18-1100"},{"key":"9250_CR210","unstructured":"Wang Y, Shi L, Wei L, Zhu W, Chen J, Wang Z, Wen S, Chen W, Wang Y, Jia J (2018c) The Sogou-TIIC speech translation system for IWSLT 2018. In: Proceedings of the 2018 international workshop on spoken language translation (IWSLT), Bruges, pp 112\u2013117"},{"key":"9250_CR207","doi-asserted-by":"crossref","unstructured":"Wang C, Wu Y, Liu S, Yang Z, Zhou M (2019a) Bridging the gap between pre-training and fine-tuning for end-to-end speech translation. Computing research repository arXiv:1909.07575","DOI":"10.1609\/aaai.v34i05.6452"},{"key":"9250_CR209","doi-asserted-by":"crossref","unstructured":"Wang X, Wu J, Chen J, Li L, Wang Y, Wang WY (2019b) VATEX: a large-scale, high-quality multilingual dataset for video-and-language research. Computing research repository arXiv:1904.03493","DOI":"10.1109\/ICCV.2019.00468"},{"key":"9250_CR211","doi-asserted-by":"crossref","unstructured":"Watanabe S, Hori T, Karita S, Hayashi T, Nishitoba J, Unno Y, Soplin NEY, Heymann J, Wiesner M, Chen N, et\u00a0al (2018) ESPnet: end-to-end speech processing toolkit. In: Proceedings of Interspeech, Hyderabad, pp 2207\u20132211","DOI":"10.21437\/Interspeech.2018-1456"},{"key":"9250_CR212","doi-asserted-by":"crossref","unstructured":"Weiss RJ, Chorowski J, Jaitly N, Wu Y, Chen Z (2017) Sequence-to-sequence models can directly translate foreign speech. In: Proceedings of Interspeech, Stockholm","DOI":"10.21437\/Interspeech.2017-503"},{"key":"9250_CR213","unstructured":"Wu Z, Caglayan O, Ive J, Wang J, Specia L (2019a) Transformer-based cascaded multimodal speech translation. In: Proceedings of the 16th international workshop on spoken language translation (IWSLT), Hong Kong"},{"key":"9250_CR214","unstructured":"Wu Z, Ive J, Wang J, Madhyastha P, Specia L (2019b) Predicting actions to help predict translations. In: Proceedings of the how2 challenge: new tasks for vision and language, Long Beach"},{"key":"9250_CR215","doi-asserted-by":"crossref","unstructured":"Xiao J, Hays J, Ehinger KA, Oliva A, Torralba A (2010) SUN database: Large-scale scene recognition from abbey to zoo. The 23rd IEEE conference on computer vision and pattern recognition. CVPR, San Francisco, pp 3485\u20133492","DOI":"10.1109\/CVPR.2010.5539970"},{"key":"9250_CR216","unstructured":"Xu K, Ba J, Kiros R, Cho K, Courville A, Salakhudinov R, Zemel R, Bengio Y (2015) Show, attend and tell: neural image caption generation with visual attention. In: Proceedings of the 32nd international conference on machine learning (ICML), JMLR workshop and conference proceedings, Lille, pp 2048\u20132057"},{"key":"9250_CR217","doi-asserted-by":"crossref","unstructured":"Yao B, Jiang X, Khosla A, Lin AL, Guibas L, Fei-Fei L (2011) Human action recognition by learning bases of action attributes and parts. In: Proceedings of the 2011 IEEE international conference on computer vision (ICCV), Barcelona, pp 1331\u20131338","DOI":"10.1109\/ICCV.2011.6126386"},{"key":"9250_CR218","doi-asserted-by":"crossref","unstructured":"Yoshikawa Y, Shigeto Y, Takeuchi A (2017) STAIR captions: Constructing a large-scale Japanese image caption dataset. In: Proceedings of the 55th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Vancouver, pp 417\u2013421","DOI":"10.18653\/v1\/P17-2066"},{"key":"9250_CR219","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1162\/tacl_a_00166","volume":"2","author":"P Young","year":"2014","unstructured":"Young P, Lai A, Hodosh M, Hockenmaier J (2014) From image descriptions to visual denotations: new similarity metrics for semantic inference over event descriptions. Trans Assoc Comput Linguistics 2:67\u201378","journal-title":"Trans Assoc Comput Linguistics"},{"key":"9250_CR220","volume-title":"Automatic speech recognition: a deep learning approach","author":"D Yu","year":"2016","unstructured":"Yu D, Deng L (2016) Automatic speech recognition: a deep learning approach. Springer, Berlin"},{"issue":"3","key":"9250_CR221","doi-asserted-by":"publisher","first-page":"396","DOI":"10.1109\/JAS.2017.7510508","volume":"4","author":"D Yu","year":"2017","unstructured":"Yu D, Li J (2017) Recent progresses in deep learning based acoustic models. IEEE\/CAA J Autom Sin 4(3):396\u2013409","journal-title":"IEEE\/CAA J Autom Sin"},{"key":"9250_CR222","unstructured":"Zadeh A, Zellers R, Pincus E, Morency LP (2016) MOSI: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos. Computing research repository arXiv:1606.06259"},{"key":"9250_CR223","doi-asserted-by":"crossref","unstructured":"Zhang J, Utiyama M, Sumita E, Neubig G, Nakamura S (2017) NICT-NAIST system for WMT17 multimodal translation task. In: Proceedings of the 2nd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Copenhagen, pp 477\u2013482","DOI":"10.18653\/v1\/W17-4753"},{"key":"9250_CR224","doi-asserted-by":"crossref","unstructured":"Zhang P, Ge N, Chen B, Fan K (2019) Lattice transformer for speech translation. In: Proceedings of the 57th annual meeting of the association for computational linguistics (ACL), Association for Computational Linguistics (ACL), Florence, pp 6475\u20136484","DOI":"10.18653\/v1\/P19-1649"},{"key":"9250_CR225","doi-asserted-by":"crossref","unstructured":"Zheng R, Yang Y, Ma M, Huang L (2018) Ensemble sequence level training for multimodal MT: OSU-Baidu WMT18 multimodal machine translation system report. In: Proceedings of the 3rd conference on machine translation (WMT), Association for Computational Linguistics (ACL), Belgium, pp 638\u2013642","DOI":"10.18653\/v1\/W18-6443"},{"issue":"5","key":"9250_CR226","doi-asserted-by":"publisher","first-page":"1180","DOI":"10.1109\/JPROC.2013.2249491","volume":"101","author":"B Zhou","year":"2013","unstructured":"Zhou B (2013) Statistical machine translation for speech: a perspective on structures, learning, and decoding. Proc IEEE 101(5):1180\u20131202","journal-title":"Proc IEEE"},{"key":"9250_CR227","doi-asserted-by":"crossref","unstructured":"Zhou M, Cheng R, Lee YJ, Yu Z (2018) A visual attention grounding neural model for multimodal machine translation. In: Proceedings of the 2018 conference on empirical methods in natural language processing (EMNLP), Association for Computational Linguistics (ACL), Brussels, pp 3643\u20133653","DOI":"10.18653\/v1\/D18-1400"}],"container-title":["Machine Translation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10590-020-09250-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10590-020-09250-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10590-020-09250-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,11]],"date-time":"2024-08-11T19:31:08Z","timestamp":1723404668000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10590-020-09250-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,8,13]]},"references-count":227,"journal-issue":{"issue":"2-3","published-print":{"date-parts":[[2020,9]]}},"alternative-id":["9250"],"URL":"https:\/\/doi.org\/10.1007\/s10590-020-09250-0","relation":{},"ISSN":["0922-6567","1573-0573"],"issn-type":[{"value":"0922-6567","type":"print"},{"value":"1573-0573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,8,13]]},"assertion":[{"value":"5 December 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 July 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 August 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}