{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:46:09Z","timestamp":1760240769267,"version":"build-2065373602"},"reference-count":29,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2019,9,1]],"date-time":"2019-09-01T00:00:00Z","timestamp":1567296000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Informatics"],"abstract":"<jats:p>To build state-of-the-art Neural Machine Translation (NMT) systems, high-quality parallel sentences are needed. Typically, large amounts of data are scraped from multilingual web sites and aligned into datasets for training. Many tools exist for automatic alignment of such datasets. However, the quality of the resulting aligned corpus can be disappointing. In this paper, we present a tool for automatic misalignment detection (MAD). We treated the task of determining whether a pair of aligned sentences constitutes a genuine translation as a supervised regression problem. We trained our algorithm on a manually labeled dataset in the FR\u2013NL language pair. Our algorithm used shallow features and features obtained after an initial translation step. We showed that both the Levenshtein distance between the target and the translated source, as well as the cosine distance between sentence embeddings of the source and the target were the two most important features for the task of misalignment detection. Using gold standards for alignment, we demonstrated that our model can increase the quality of alignments in a corpus substantially, reaching a precision close to 100%. Finally, we used our tool to investigate the effect of misalignments on NMT performance.<\/jats:p>","DOI":"10.3390\/informatics6030035","type":"journal-article","created":{"date-parts":[[2019,9,2]],"date-time":"2019-09-02T03:16:12Z","timestamp":1567394172000},"page":"35","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Misalignment Detection for Web-Scraped Corpora: A Supervised Regression Approach"],"prefix":"10.3390","volume":"6","author":[{"given":"Arne","family":"Defauw","sequence":"first","affiliation":[{"name":"CrossLang NV, 9050 Gentbrugge, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sara","family":"Szoc","sequence":"additional","affiliation":[{"name":"CrossLang NV, 9050 Gentbrugge, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anna","family":"Bardadym","sequence":"additional","affiliation":[{"name":"CrossLang NV, 9050 Gentbrugge, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joris","family":"Brabers","sequence":"additional","affiliation":[{"name":"CrossLang NV, 9050 Gentbrugge, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Frederic","family":"Everaert","sequence":"additional","affiliation":[{"name":"CrossLang NV, 9050 Gentbrugge, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Roko","family":"Mijic","sequence":"additional","affiliation":[{"name":"Independent Data Science Consultant, 9000 Ghent, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kim","family":"Scholte","sequence":"additional","affiliation":[{"name":"CrossLang NV, 9050 Gentbrugge, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tom","family":"Vanallemeersch","sequence":"additional","affiliation":[{"name":"CrossLang NV, 9050 Gentbrugge, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Koen","family":"Van Winckel","sequence":"additional","affiliation":[{"name":"CrossLang NV, 9050 Gentbrugge, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joachim","family":"Van den Bogaert","sequence":"additional","affiliation":[{"name":"CrossLang NV, 9050 Gentbrugge, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,9,1]]},"reference":[{"key":"ref_1","unstructured":"Goutte, C., Carpaut, M., and Foster, G. (November, January 28). The impact of Sentence Alignment Errors on Phrase-Based Machine Translation Performance. Proceedings of the Tenth Biennial Conference of the Association for Machine Translation in the Americas, San Diego, CA, USA."},{"key":"ref_2","unstructured":"Chen, B., Kuhn, R., Foster, G., Cherry, C., and Huang, F. (November, January 29). Bilingual Methods for Adaptive Training Data Selection for Machine Translation. Proceedings of the 12th Conference of the Association for Machine Translation in the Americas (AMTA), Austin, TX, USA."},{"key":"ref_3","unstructured":"Belinkov, Y., and Bisk, Y. (2016). Synthetic and Natural Noise Both Break Neural Machine Translation. CoRR."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Khayrallah, H., and Koehn, P. (2018). On the Impact of Various Types of Noise on Neural Machine Translation. arXiv.","DOI":"10.18653\/v1\/W18-2709"},{"key":"ref_5","unstructured":"Lamraoui, F., and Langlais, P. (2013, January 2\u20136). Yet Another Fast and Open Source Sentence Aligner. Time to Reconsider Sentence Alignment?. Proceedings of the Machine Translation Summit XIV, Nice, France."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"477","DOI":"10.1162\/089120105775299168","article-title":"Improving Machine Translation Performance by Exploiting Non-Parallel Corpora","volume":"31","author":"Munteanu","year":"2005","journal-title":"Comput. Linguist."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Etchegoyhen, T., and Azpeitia, A. (2016, January 7\u201312,). Set-Theoretic Alignment for Comparable Corpora. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, Berlin, Germany.","DOI":"10.18653\/v1\/P16-1189"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Abdul-Rauf, S., and Schwenk, H. (April, January 30). On the Use of Comparable Corpora to Improve SMT performance. Proceedings of the 12th Conference of the European Chapter of the ACL (EACL 2009), Athens, Greece.","DOI":"10.3115\/1609067.1609068"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1340","DOI":"10.1109\/JSTSP.2017.2764273","article-title":"An Empirical Analysis of NMT-Derived Interlingual Embeddings and their Use in Parallel Sentence Identification","volume":"11","year":"2017","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Gr\u00e9goire, F., and Langlais, P. (2017, January 3). A First Attempt Toward a Deep Learning Framework for Identifying Parallel Sentences in Comparable Corpora. Proceedings of the 10th Workshop on Building and Using Comparable Corpora, Vancouver, Canada.","DOI":"10.18653\/v1\/W17-2509"},{"key":"ref_11","unstructured":"Carpuat, M., Vyas, Y., and Niu, X. (August, January 30). Detecting Cross-lingual Semantic Divergence for Neural Machine Translation. Proceedings of the First Workshop on Neural Machine Translation, Vancouver, Canada."},{"key":"ref_12","unstructured":"Gr\u00e9goire, F., and Langlais, P. (2018, January 21\u201325). Extracting Parallel Sentences with Bidirectional Recurrent Neural Networks to Improve Machine Translation. Proceedings of the 27th International Conference on Computational Linguistics, Santa Fe, NM, USA."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Schwenk, H. (2018, January 15\u201320). Filtering and Mining Parallel Data in a Joint Multilingual Space. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia.","DOI":"10.18653\/v1\/P18-2037"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Vyas, Y., Niu, X., and Carpuat, M. (2018, January 1\u20136). Identifying Semantic Divergences in Parallel Text without Annotations. Proceedings of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies ACL, New Orleans, LA, USA.","DOI":"10.18653\/v1\/N18-1136"},{"key":"ref_15","unstructured":"Bouamor, H., and Sajjad, H. (2018, January 7\u201312). Parallel Sentence Extraction from Comparable Corpora using Multilingual Sentence Embeddings. Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC), Miyazaki, Japan."},{"key":"ref_16","unstructured":"Hassan, H., Aue, A., Chen, C., Chowdhary, V., Clark, J., Federmann, C., Huang, X., Junczys-Dowmunt, M., Lewis, W., and Li, M. (2018). Achieving Human Parity on Automatic Chinese to English Translation. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Pham, M., Crego, J., Senellart, J., and Yvon, F. (November, January 31). Fixing Translation Divergences in Parallel Corpora for Neural MT. Proceedings of the 2018 Conference on Emperical Methods in Natural Language Processing, Brussels, Belgium.","DOI":"10.18653\/v1\/D18-1328"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Artetxe, M., and Schwenk, H. (2018). Margin-Based Parallel Corpus Mining with Multilingual Sentence Embeddings. arXiv.","DOI":"10.18653\/v1\/P19-1309"},{"key":"ref_19","unstructured":"Guo, M., Shen, Q., Yang, Y., Ge, H., Cer, D., Hernandez Abrego, G., Stevens, K., Constant, N., Sung, Y., and Strope, B. (November, January 31). Effective Parallel Corpus Mining using Bilingual Sentence Embeddings. Proceedings of the Third Conference on Machine Translation (WMT), Brussels, Belgium."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"S\u00e1nchez-Cartagena, V.M., Ba\u00f1\u00f3n, M., Sergio Ortiz-Rojas, S., and Ram\u00edrez-S\u00e1nchez, G. (November, January 31). Prompsit\u2019s Submission to WMT 2018 Parallel Corpus Filtering Shared Task. Proceedings of the Third Conference on Machine Translation (WMT), Brussels, Belgium.","DOI":"10.18653\/v1\/W18-6488"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Conneau, A., Kiela, D., Schwenk, H., Barrault, L., and Bordes, A. (2017, January 7\u201311). Supervised Learning of Universal Sentence Representations from Natural Language Inference Data. Proceedings of the 2017 Conference on Emperical Methods in Natural Language Processing, Copenhagen, Denmark.","DOI":"10.18653\/v1\/D17-1070"},{"key":"ref_22","first-page":"2825","article-title":"Scikit-learn: Machine Learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Klein, G., Kim, Y., Deng, Y., Senellart, J., and Rush, A.M. (2017). OpenNMT: Open-Source Toolkit for Neural Machine Translation. arXiv.","DOI":"10.18653\/v1\/P17-4012"},{"key":"ref_24","unstructured":"Varga, D., N\u00e9meth, L., Hal\u00e1csy, P., Kornai, A., Tr\u00f3n, V., and Nagy, V. (2005, January 21\u201323). Parallel Corpora for Medium Density Languages. Proceedings of the RANLP, Borovets, Bulgaria."},{"key":"ref_25","unstructured":"Defauw, A., Vanallemeersch, T., Szoc, S., Everaert, F., Van Winckel, K., Scholte, K., Brabers, J., and Van den Bogaert, J. (2019, January 19\u201323). Collecting Domain Specific Data for MT: An Evaluation of the ParaCrawl Pipeline. Proceedings of the Machine Translation Summit, Dublin, Ireland."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Brown, P.F., Lai, J.C., and Mercer, R.L. (1991, January 18\u201321). Aligning Sentences in Parallel Corpora. Proceedings of the 29th Annual Meeting of the ACL, Berkeley, CA, USA.","DOI":"10.3115\/981344.981366"},{"key":"ref_27","first-page":"145","article-title":"The First Automatic Translation Memory Cleaning Shared Task","volume":"30","author":"Barbu","year":"2016","journal-title":"Comput. Transl."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Koehn, P., Khayrallah, H., Heafield, K., and Forcada, M.L. (November, January 31). Findings of the WMT 2018 Shared Task on Parallel Corpus Filtering. Proceedings of the Third Conference on Machine Translation (WMT), Brussels, Belgium.","DOI":"10.18653\/v1\/W18-6453"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zweigenbaum, P., Sharoff, S., and Rapp, R. (2017, January 3). Overview of the Second BUCC Shared Task: Spotting Parallel Sentences in Comparable Corpora. Proceedings of the 10th Workshop on Building and Using Comparable Corpora, Vancouver, Canada.","DOI":"10.18653\/v1\/W17-2512"}],"container-title":["Informatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2227-9709\/6\/3\/35\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:15:45Z","timestamp":1760188545000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2227-9709\/6\/3\/35"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,9,1]]},"references-count":29,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2019,9]]}},"alternative-id":["informatics6030035"],"URL":"https:\/\/doi.org\/10.3390\/informatics6030035","relation":{},"ISSN":["2227-9709"],"issn-type":[{"type":"electronic","value":"2227-9709"}],"subject":[],"published":{"date-parts":[[2019,9,1]]}}}