{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T12:55:03Z","timestamp":1760100903650},"reference-count":26,"publisher":"Cambridge University Press (CUP)","issue":"4","license":[{"start":{"date-parts":[[2016,6,15]],"date-time":"2016-06-15T00:00:00Z","timestamp":1465948800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2016,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In recent decades, statistical approaches have significantly advanced the development of machine translation systems. However, the applicability of these methods directly depends on the availability of very large quantities of parallel data. Recent works have demonstrated that a comparable corpus can compensate for the shortage of parallel corpora. In this paper, we propose an alternative to comparable corpora containing text documents as resources for extracting parallel data: a multimodal comparable corpus with audio documents in source language and text document in target language, built from<jats:italic>Euronews<\/jats:italic>and<jats:italic>TED<\/jats:italic>web sites. The audio is transcribed by an automatic speech recognition system, and translated with a baseline statistical machine translation system. We then use information retrieval in a large text corpus in the target language in order to extract parallel sentences\/phrases. We evaluate the quality of the extracted data on an English to French translation task and show significant improvements over a state-of-the-art baseline.<\/jats:p>","DOI":"10.1017\/s1351324916000152","type":"journal-article","created":{"date-parts":[[2016,6,15]],"date-time":"2016-06-15T18:25:18Z","timestamp":1466015118000},"page":"603-625","source":"Crossref","is-referenced-by-count":5,"title":["Building and using multimodal comparable corpora for machine translation"],"prefix":"10.1017","volume":"22","author":[{"given":"HAITHEM","family":"AFLI","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"LO\u00cfC","family":"BARRAULT","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"HOLGER","family":"SCHWENK","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2016,6,15]]},"reference":[{"key":"S1351324916000152_ref12","doi-asserted-by":"crossref","unstructured":"Munteanu D. S. and Marcu D. 2006. Extracting parallel sub-sentential fragments from non-parallel corpora. In Proceedings of the 21st International Conference on Computational Linguistics and the 44th Annual Meeting of the Association for Computational Linguistics, ACL-44. Sydney, Australia, pp. 81\u20138.","DOI":"10.3115\/1220175.1220186"},{"key":"S1351324916000152_ref22","unstructured":"Snover S. , Dorr B. , Schwartz R. , Micciulla M. , and Makhoul J. 2006. A study of translation edit rate with targeted human annotation. In Proceedings of Association for Machine Translation in the Americas, pp. 223\u201331."},{"key":"S1351324916000152_ref2","first-page":"263","article-title":"The mathematics of statistical machine translation: parameter estimation","volume":"19","author":"Brown","year":"1993","journal-title":"Computational Linguistics"},{"key":"S1351324916000152_ref19","unstructured":"Rousseau A. , Bougares F. , Del\u00e9glise P. , Schwenk H. , and Est\u00e8ve Y. 2011. LIUM's systems for the IWSLT 2011 speech translation tasks. In International Workshop on Spoken Language Translation 2011, San Francisco, USA."},{"key":"S1351324916000152_ref17","doi-asserted-by":"publisher","DOI":"10.1162\/089120103322711578"},{"key":"S1351324916000152_ref14","unstructured":"Papineni K. , Roukos S. , Ward T. and Zhu W.-J. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, ACL '02. Philadelphia, USA, pp. 311\u201318."},{"key":"S1351324916000152_ref11","doi-asserted-by":"publisher","DOI":"10.1162\/089120105775299168"},{"key":"S1351324916000152_ref24","doi-asserted-by":"crossref","unstructured":"Utiyama M. and Isahara H. 2003. Reliable measures for aligning japanese-english news articles and sentences. In Proceedings of the 41st Annual Meeting on Association for Computational Linguistics - Volume 1, ACL '03, pp. 72\u20139.","DOI":"10.3115\/1075096.1075106"},{"key":"S1351324916000152_ref18","unstructured":"Riesa J. and Marcu D. 2012. Automatic parallel fragment extraction from noisy data. In Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL HLT '12. Montreal, Quebec, Canada, pp. 538\u201342."},{"key":"S1351324916000152_ref20","unstructured":"Rousseau A. , Del\u00e9glise P. and Est\u00e8ve Y. 2012. Ted-lium: an automatic speech recognition dedicated corpus. In Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12), Istanbul, Turkey."},{"key":"S1351324916000152_ref26","unstructured":"Zhao B. and Vogel S. 2002. Adaptive parallel sentences mining from web bilingual news collection. In Proceedings of the 2002 IEEE International Conference on Data Mining, ICDM '02, Washington, DC, USA, IEEE Computer Society."},{"key":"S1351324916000152_ref8","unstructured":"Hewavitharana S. and Vogel S. 2011. Extracting parallel phrases from comparable data. In Proceedings of the 4th Workshop on Building and Using Comparable Corpora: Comparable Corpora and the Web, BUCC '11, Portland, Oregon, USA, pp. 61\u20138."},{"key":"S1351324916000152_ref4","doi-asserted-by":"crossref","unstructured":"Del\u00e9glise P. , Est\u00e8ve Y. , Meignier S. and Merlin T. 2009. Improvements to the LIUM french ASR system based on CMU Sphinx: what helps to significantly reduce the word error rate? In Interspeech 2009, Brighton, UK.","DOI":"10.21437\/Interspeech.2009-607"},{"key":"S1351324916000152_ref25","doi-asserted-by":"publisher","DOI":"10.1002\/asi.10261"},{"key":"S1351324916000152_ref5","doi-asserted-by":"crossref","unstructured":"Fung P. and Cheung P. 2004. Multi-level bootstrapping for extracting parallel sentences from a quasi-comparable corpus. In Proceedings of the 20th International Conference on Computational Linguistics, COLING '04. Geneva, Switzerland.","DOI":"10.3115\/1220355.1220506"},{"key":"S1351324916000152_ref1","doi-asserted-by":"publisher","DOI":"10.1007\/s10590-011-9114-9"},{"key":"S1351324916000152_ref23","unstructured":"Stolcke A. 2002. SRILM - an extensible language modeling toolkit. In Proceedings of the International Conference on Spoken Language Processing, pp. 257\u201386."},{"key":"S1351324916000152_ref9","doi-asserted-by":"crossref","unstructured":"Koehn P. , Hoang H. , Birch A. , Callison-Burch C. , Federico M. , Bertoldi N. , Cowan B. , Shen W. , Moran C. , Zens R. , Dyer C. , Bojar O. , Constantin A. , and Herbst E. 2007. Moses: open source toolkit for statistical machine translation. In Proceedings of the 45th Annual Meeting of the ACL on Interactive Poster and Demonstration Sessions, ACL '07. Prague, Czech Republic, pp. 177\u201380.","DOI":"10.3115\/1557769.1557821"},{"key":"S1351324916000152_ref10","doi-asserted-by":"crossref","unstructured":"Koehn P. , Och F. J. and Marcu D. 2003. Statistical phrase-based translation. In Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology - Volume 1, NAACL '03. Edmonton, Canada, pp. 48\u201354.","DOI":"10.21236\/ADA461156"},{"key":"S1351324916000152_ref16","unstructured":"Quirk Q. , Udupa R. and Menezes A. 2007. Generative models of noisy translations with applications to parallel fragment extraction. In In Proceedings of MT Summit XI, European Association for Machine Translation, Copenhagen, Denmark."},{"key":"S1351324916000152_ref3","unstructured":"Cettolo M. , Federico M. and Bertoldi N. 2010. Mining parallel fragments from comparable texts. In Proceedings of the 7th International Workshop on Spoken Language Translation, Paris, France."},{"key":"S1351324916000152_ref15","doi-asserted-by":"crossref","unstructured":"Paulik M. and Waibel A. 2009. Automatic translation from parallel speech: simultaneous interpretation as mt training data. ASRU, Merano, Italy.","DOI":"10.1109\/ASRU.2009.5372880"},{"key":"S1351324916000152_ref7","unstructured":"Gr\u00e9zl F. and Fousek P. 2008. Optimizing bottle-neck features for LVCSR. In Proceedings of the 2008 IEEE International Conference on Acoustics, Speech, and Signal Processing, IEEE Signal Processing Society, Las Vegas, USA, pp. 4729\u201332."},{"key":"S1351324916000152_ref6","doi-asserted-by":"crossref","unstructured":"Gao Q. and Vogel S. 2008. Parallel implementations of word alignment tool. In Software Engineering, Testing, and Quality Assurance for Natural Language Processing, SETQA-NLP '08, Columbus, Ohio, USA, pp. 49\u201357.","DOI":"10.3115\/1622110.1622119"},{"key":"S1351324916000152_ref13","doi-asserted-by":"crossref","unstructured":"Ogilvie P. and Callan J. 2001. Experiments using the lemur toolkit. In Procedding of the Trenth Text Retrieval Conference (TREC-10). National Institute of Standards and Technology Special Publication 500-207.","DOI":"10.6028\/NIST.SP.500-250.cmu-lti"},{"key":"S1351324916000152_ref21","unstructured":"Schwenk H. 2008. Investigations on large-scale lightly-supervised training for statistical machine translation. In Proceedings of the International Workshop on Spoken Language Translation. Waikiki, Hawai'i, USA, pp. 182\u201389."}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324916000152","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,17]],"date-time":"2024-06-17T17:16:34Z","timestamp":1718644594000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324916000152\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,6,15]]},"references-count":26,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2016,7]]}},"alternative-id":["S1351324916000152"],"URL":"https:\/\/doi.org\/10.1017\/s1351324916000152","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,6,15]]}}}