{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T20:50:49Z","timestamp":1773521449574,"version":"3.50.1"},"reference-count":22,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2020,11,23]],"date-time":"2020-11-23T00:00:00Z","timestamp":1606089600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,11,23]],"date-time":"2020-11-23T00:00:00Z","timestamp":1606089600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100012520","name":"Sciences et Technologies de l\u2019information et de la Communication, Universit\u00e9 Paris-Saclay","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100012520","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100000848","name":"University of Edinburgh","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100000848","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Lang Resources &amp; Evaluation"],"published-print":{"date-parts":[[2021,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We present a new English\u2013French dataset for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue. The test set contains 144 spontaneous dialogues (5700+ sentences) between native English and French speakers, mediated by one of two neural MT systems in a range of role-play settings. The dialogues are accompanied by fine-grained sentence-level judgments of MT quality, produced by the dialogue participants themselves, as well as by manually normalised versions and reference translations produced<jats:italic>a posteriori<\/jats:italic>. The motivation for the corpus is twofold:\u00a0to provide (i)\u00a0a unique resource for evaluating MT models, and (ii)\u00a0a corpus for the analysis of MT-mediated communication. We provide an initial analysis of the corpus to confirm that the participants\u2019 judgments reveal perceptible differences in MT quality between the two MT systems used.<\/jats:p>","DOI":"10.1007\/s10579-020-09514-4","type":"journal-article","created":{"date-parts":[[2020,11,23]],"date-time":"2020-11-23T17:14:43Z","timestamp":1606151683000},"page":"635-660","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["DiaBLa: a corpus of bilingual spontaneous written dialogues for machine translation"],"prefix":"10.1007","volume":"55","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9553-1768","authenticated-orcid":false,"given":"Rachel","family":"Bawden","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eric","family":"Bilinski","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thomas","family":"Lavergne","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sophie","family":"Rosset","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,11,23]]},"reference":[{"key":"9514_CR1","unstructured":"Bahdanau, D., Cho, K., & Bengio, Y. (2015). Neural machine translation by jointly learning to align and translate. In Proceedings of the 3rd international conference on learning representations, ICLR\u201915."},{"key":"9514_CR2","unstructured":"Bawden, R. (2018). Going beyond the sentence: Contextual Machine Translation of Dialogue, Ph.D. thesis. Saint-Aubin: Universit\u00e9 Paris-Saclay."},{"key":"9514_CR3","unstructured":"Bawden, R., Sennrich, R., Birch, A., & Haddow, B. (2018). Evaluating l. In Proceedings of the 2018 conference of the North American chapter of the association for computational linguistics; human language technologies, New Orleans, Louisiana, USA, NAACL-HLT\u201918 (pp. 1304\u20131313)."},{"key":"9514_CR4","unstructured":"Dyer, C., Chahuneau, V., & Smith, N. A. (2013). A simple, fast, and effective reparameterization of IBM Model 2. In Proceedings of the conference of the North American chapter of the association for computational linguistics: Human language technologies, Denver, Colorado, USA, NAACL-HLT\u201913 (pp. 644\u2013648)."},{"key":"9514_CR5","volume-title":"Microsoft speech language translation (MSLT) corpus: The IWSLT 2016 release for English","author":"C Federmann","year":"2016","unstructured":"Federmann, C., & Lewis, W. (2016). Microsoft speech language translation (MSLT) corpus: The IWSLT 2016 release for English. Redmond: Microsoft Research."},{"issue":"1","key":"9514_CR6","doi-asserted-by":"publisher","first-page":"87","DOI":"10.2307\/2340521","volume":"85","author":"RA Fisher","year":"1922","unstructured":"Fisher, R. A. (1922). On the interpretation of $$\\chi ^2$$ from contingency tables, and the calculation of P. Journal of the Royal Statistical Society, 85(1), 87\u201394.","journal-title":"Journal of the Royal Statistical Society"},{"key":"9514_CR7","unstructured":"Higashinaka, R., Funakoshi, K., Kobayashi, Y., & Inaba, M. (2016) The dialogue breakdown detection challenge: Task description, datasets, and evaluation metrics. In Proceedings of the 10th international conference on language resources and evaluation, Portoro\u017e, Slovenia, LREC\u201916 (pp. 3146\u20133150)."},{"key":"9514_CR8","doi-asserted-by":"crossref","unstructured":"Isabelle, P., Cherry, C., & Foster, G. (2017). A challenge set approach to evaluating machine translation. In Proceedings of the 2017 conference on empirical methods in natural language processing, Copenhagen, Denmark, EMNLP\u201917 (pp. 2476\u20132486).","DOI":"10.18653\/v1\/D17-1263"},{"key":"9514_CR9","doi-asserted-by":"crossref","unstructured":"Junczys-Dowmunt, M., Grundkiewicz, R., Dwojak, T., Hoang, H., Heafield, K., & Neckermann, T., et al. (2018) Marian: Fast neural machine translation in C++. arXiv:180400344 [cs] ArXiv:1804.00344.","DOI":"10.18653\/v1\/P18-4020"},{"key":"9514_CR10","doi-asserted-by":"crossref","unstructured":"King, M., & Falkedal, K. (1990). Using test suites in evaluation of machine translation systems. In Proceedings of the 1990 conference on computational linguistics, Helsinki, Finland, COLING\u201990 (pp. 211\u2013216).","DOI":"10.3115\/997939.997976"},{"key":"9514_CR11","doi-asserted-by":"crossref","unstructured":"Koehn, P., Hoang, H., Birch, A., Callison-Burch, C., Federico, & M., Bertoldi, N., et al. (2007). Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th annual meeting of the association for computational linguistics, Prague, Czech Republic, ACL\u201907 (pp. 177\u2013180).","DOI":"10.3115\/1557769.1557821"},{"key":"9514_CR12","unstructured":"Lison, P., & Tiedemann, J. (2016). OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles. In Proceedings of the 10th language resources and evaluation conference, Portoro\u017e, Slovenia, LREC\u201916 (pp. 923\u2013929)"},{"key":"9514_CR13","unstructured":"Lison, P., Tiedemann, J., & Kouylekov, M. (2018). OpenSubtitles2018: Statistical rescoring of sentence alignments in large, noisy parallel corpora. In Proceedings of the eleventh international conference on language resources and evaluation (LREC), 2018. Miyazaki: European Language Resources Association (ELRA)."},{"key":"9514_CR14","doi-asserted-by":"crossref","unstructured":"Morimoto, T., Uratani, N., Takezawa, T., Furuse, O., Sobashima, Y., & Iida, H., et al. (1994). A speech and language database for speech translation research. In Proceedings of the 3rd international conference on spoken language processing, Yokohama, Japan, ICSLP\u201994 (pp. 1791\u20131794).","DOI":"10.21437\/ICSLP.1994-450"},{"key":"9514_CR15","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the association for computational linguistics, Philadelphia, Pennsylvania, USA, ACL\u201902 (pp. 311\u2013318).","DOI":"10.3115\/1073083.1073135"},{"key":"9514_CR18","doi-asserted-by":"crossref","unstructured":"Sennrich, R., Firat, O., Cho, K., Birch, A., Haddow, B., & Hitschler, J., et al. (2017). Nematus: A toolkit for neural machine translation. In Proceedings of the software demonstrations of the 15th conference of the European chapter of the association for computational linguistics, Valencia, Spain, EACL\u201917 (pp. 65\u201368).","DOI":"10.18653\/v1\/E17-3017"},{"key":"9514_CR16","doi-asserted-by":"crossref","unstructured":"Sennrich, R., Haddow, B., & Birch, A. (2016a). Controlling politeness in neural machine translation via side constraints. In Proceedings of the 2016 conference of the North American chapter of the association for computational linguistics: Human language technologies, San Diego, California, USA, NAACL-HLT\u201916 (pp. 35\u201340).","DOI":"10.18653\/v1\/N16-1005"},{"key":"9514_CR17","doi-asserted-by":"crossref","unstructured":"Sennrich, R., Haddow, B., & Birch, A. (2016b). Neural machine translation of rare words with subword units. In Proceedings of the 54th annual meeting of the association for computational linguistics, Berlin, Germany, ACL\u201916 (pp. 1715\u20131725).","DOI":"10.18653\/v1\/P16-1162"},{"issue":"3","key":"9514_CR20","first-page":"303","volume":"12","author":"T Takezawa","year":"2007","unstructured":"Takezawa, T., Kikui, G., Mizushima, M., & Sumita, E. (2007). Multilingual spoken language corpus development for communication research. Computational Linguistics and Chinese Language Processing, 12(3), 303\u2013324.","journal-title":"Computational Linguistics and Chinese Language Processing"},{"key":"9514_CR19","unstructured":"Takezawa, T., Sumita, E., Sugaya, F., Yamamoto, H., & Yamamoto, S. (2002). Toward a broad-coverage bilingual corpus for speech translation of travel conversations in the real world. In Proceedings of the 3rd international conference on language resources and evaluation, Las Palmas, Canary Islands, Spain, LREC\u201902 (pp. 147\u2013152)."},{"key":"9514_CR21","doi-asserted-by":"crossref","unstructured":"Tiedemann, J., & Scherrer, Y. (2017). Neural machine translation with extended context. In Proceedings of the 3rd workshop on discourse in machine translation, Copenhagen, Denmark, DISCOMT\u201917 (pp. 82\u201392).","DOI":"10.18653\/v1\/W17-4811"},{"key":"9514_CR22","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/978-3-662-04230-4_1","volume-title":"Verbmobil: Foundations of speech-to-speech translation","author":"W Wahlster","year":"2000","unstructured":"Wahlster, W. (2000). Mobile speech-to-speech translation of spontaneous dialogs: An overview of the final Verbmobil system. Verbmobil: Foundations of speech-to-speech translation (pp. 3\u201321). Berlin, Heidelberg: Springer."}],"container-title":["Language Resources and Evaluation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10579-020-09514-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10579-020-09514-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10579-020-09514-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,29]],"date-time":"2022-11-29T10:26:52Z","timestamp":1669717612000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10579-020-09514-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,23]]},"references-count":22,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,9]]}},"alternative-id":["9514"],"URL":"https:\/\/doi.org\/10.1007\/s10579-020-09514-4","relation":{},"ISSN":["1574-020X","1574-0218"],"issn-type":[{"value":"1574-020X","type":"print"},{"value":"1574-0218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,11,23]]},"assertion":[{"value":"24 October 2020","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 November 2020","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}