{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,29]],"date-time":"2025-11-29T07:55:10Z","timestamp":1764402910426,"version":"build-2065373602"},"reference-count":51,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2019,1,18]],"date-time":"2019-01-18T00:00:00Z","timestamp":1547769600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Machine translation is used in many applications in everyday life. Due to the increase of translated documents that need to be organized as useful or not (for building a translation model), the automated categorization of texts (classification), is a popular research field of machine learning. This kind of information can be quite helpful for machine translation. Our parallel corpora (English-Greek and English-Italian) are based on educational data, which are quite difficult to translate. We apply two state of the art architectures, Random Forest (RF) and Deeplearnig4j (DL4J), to our data (which constitute three translation outputs). To our knowledge, this is the first time that deep learning architectures are applied to the automatic selection of parallel data. We also propose new string-based features that seem to be effective for the classifier, and we investigate whether an attribute selection method could be used for better classification accuracy. Experimental results indicate an increase of up to 4% (compared to our previous work) using RF and rather satisfactory results using DL4J.<\/jats:p>","DOI":"10.3390\/a12010026","type":"journal-article","created":{"date-parts":[[2019,1,18]],"date-time":"2019-01-18T11:26:55Z","timestamp":1547810815000},"page":"26","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Ensemble and Deep Learning for Language-Independent Automatic Selection of Parallel Data"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2844-5488","authenticated-orcid":false,"given":"Despoina","family":"Mouratidis","sequence":"first","affiliation":[{"name":"Department of Informatics, Ionian University, 491 00 Kerkira, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3270-5078","authenticated-orcid":false,"given":"Katia Lida","family":"Kermanidis","sequence":"additional","affiliation":[{"name":"Department of Informatics, Ionian University, 491 00 Kerkira, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,1,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Collobert, R., and Weston, J. (2008, January 5\u20139). A unified architecture for natural language processing: Deep neural networks with multitask learning. Proceedings of the 25th International Conference on Machine learning, Helsinki, Finland.","DOI":"10.1145\/1390156.1390177"},{"key":"ref_2","first-page":"2493","article-title":"Natural language processing (almost) from scratch","volume":"12","author":"Collobert","year":"2011","journal-title":"JMLR"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Koehn, P., Och, F.J., and Marcu, D. (June, January 27). Statistical phrase-based translation. Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology, Edmonton, AB, Canada.","DOI":"10.3115\/1073445.1073462"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Bentivogli, L., Bisazza, A., Cettolo, M., and Federico, M. (2016, January 1\u20135). Neural versus phrase-based machine translation quality: A case study. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, TX, USA.","DOI":"10.18653\/v1\/D16-1025"},{"key":"ref_5","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2015, January 7\u20139). Neural machine translation by jointly learning to align and translate. Proceedings of the 3th International Conference on Learning Representations, San Diego, CA, USA."},{"key":"ref_6","unstructured":"Peris, \u00c1., Cebri\u00e1n, L., and Casacuberta, F. (arXiv, 2017). Online Learning for Neural Machine Translation Post-editing, arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"ref_8","unstructured":"Mnih, A., and Hinton, G.E. (2009, January 8\u201311). A scalable hierarchical distributed language model. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"21","DOI":"10.5120\/8626-2492","article-title":"Comparative analysis of classification algorithms on different datasets using WEKA","volume":"54","author":"Arora","year":"2012","journal-title":"IJCA"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Mouratidis, D., and Kermanidis, K.L. (2018, January 25\u201327). Automatic Selection of Parallel Data for Machine Translation. Proceedings of the IFIP International Conference on Artificial Intelligence Applications and Innovations, Rhodes, Greece.","DOI":"10.1007\/978-3-319-92016-0_14"},{"key":"ref_11","unstructured":"Kalchbrenner, N., and Blunsom, P. (2013, January 18\u201321). Recurrent continuous translation models. Proceedings of the ACL Conference on Empirical Methods in Natural Language Processing (EMNLP), Seattle, WA, USA."},{"key":"ref_12","unstructured":"Kyunghyun, C., Bart, V.M., Caglar, G., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014, January 25\u201329). Learning phrase representations using RNN encoder-decoder for statistical machine translation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Cho, K., Van Merri\u00ebnboer, B., Bahdanau, D., and Bengio, Y. (2014, January 25). On the properties of neural machine translation: Encoder-decoder approaches. Proceedings of the SSST-8, Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation, Doha, Qatar.","DOI":"10.3115\/v1\/W14-4012"},{"key":"ref_14","unstructured":"Hill, F., Cho, K., Jean, S., Devin, C., and Bengio, Y. (arXiv, 2015). Embedding word similarity with neural machine translation, arXiv."},{"key":"ref_15","unstructured":"Sutskever, I., Vinyals, O., and Le, Q.V. (2014, January 8\u201313). Sequence to sequence learning with neural networks. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Skansi, S. (2018). Introduction to Deep Learning: From Logical Calculus to Artificial Intelligence, Springer.","DOI":"10.1007\/978-3-319-73004-2"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"440","DOI":"10.1093\/bioinformatics\/btp621","article-title":"Pitfalls of supervised feature selection","volume":"26","author":"Smialowski","year":"2009","journal-title":"Bioinformatics"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Bordes, A., Chopra, S., and Weston, J. (2014, January 25\u201329). Question answering with subgraph embeddings. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1067"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"14","DOI":"10.5120\/17456-8202","article-title":"Feature Selection based Classification using Naive Bayes, J48 and Support Vector Machine","volume":"99","author":"Bhosale","year":"2014","journal-title":"IJCA"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Qiang, G. (2010, January 7\u201310). An effective algorithm for improving the performance of Naive Bayes for text classification. Proceedings of the Second International Conference on Computer Research and Development, Kuala Lumpur, Malaysia.","DOI":"10.1109\/ICCRD.2010.160"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Mohamed, W.N.H.W., Salleh, M.N.M., and Omar, A.H. (2012, January 23\u201325). A comparative study of reduced error pruning method in decision tree algorithms. Proceedings of the IEEE International Conference on Control System, Computing and Engineering (ICCSCE), Penang, Malaysia.","DOI":"10.1109\/ICCSCE.2012.6487177"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1051\/matecconf\/20164206002","article-title":"Performance Comparison of Feature Selection Methods","volume":"42","author":"Phyu","year":"2016","journal-title":"MATEC Web Conf."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Mulay, S.A., Devale, P.R., and Garje, G.V. (2010, January 11\u201312). Decision tree based support vector machine for intrusion detection. Proceedings of the International Conference on Networking and Information Technology (ICNIT), Manila, Philippines.","DOI":"10.1109\/ICNIT.2010.5508557"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Bosch, A., Zisserman, A., and Munoz, X. (2007, January 14\u201320). Image classification using random forests and ferns. Proceedings of the IEEE 11th International Conference on Computer Vision (ICCV), Rio de Janeiro, Brazil.","DOI":"10.1109\/ICCV.2007.4409066"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1915","DOI":"10.1109\/TPAMI.2012.231","article-title":"Learning hierarchical features for scene labeling","volume":"35","author":"Farabet","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Koehn, P., Hoang, H., Birch, A., Callison-Burch, C., Federico, M., Bertoldi, N., and Dyer, C. (2007, January 25\u201327). Moses: Open source toolkit for statistical machine translation. Proceedings of the 45th Annual Meeting of the ACL on Interactive Poster and Demonstration Sessions, Prague, Czech Republic.","DOI":"10.3115\/1557769.1557821"},{"key":"ref_27","first-page":"217","article-title":"Random forest classifier for remote sensing classification","volume":"26","author":"Pal","year":"2005","journal-title":"IJRS"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2913","DOI":"10.4304\/jcp.7.12.2913-2920","article-title":"An Improved Random Forest Classifier for Text Categorization","volume":"7","author":"Xu","year":"2012","journal-title":"J. Comput."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"2999","DOI":"10.1016\/j.rse.2008.02.011","article-title":"Evaluation of Random Forest and Adaboost tree-based ensemble classification and spectral band selection for ecotope mapping using airborne hyperspectral imagery","volume":"112","author":"Chan","year":"2008","journal-title":"Remote Sens. Environ."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Assun\u00e7ao, F., Louren\u00e7o, N., Machado, P., and Ribeiro, B. (arXiv, 2018). DENSER: Deep Evolutionary Network Structured Representation, arXiv.","DOI":"10.1007\/s10710-018-9339-y"},{"key":"ref_31","unstructured":"Snoek, J., Rippel, O., Swersky, K., Kiros, R., Satish, N., Sundaram, N., Patwary, M., Prabhat, M., and Adams, R. (2015, January 7\u20139). Scalable bayesian optimization using deep neural networks. Proceedings of the 32nd International Conference on Machine Learning, Lille, France."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1109\/MSP.2012.2205597","article-title":"Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups","volume":"29","author":"Hinton","year":"2012","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_34","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). ImageNet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_35","unstructured":"Pighin, D., M\u00e0rquez, L., and May, J. (2012, January 21\u201327). An Analysis (and an Annotated Corpus) of User Responses to Machine Translation Output. Proceedings of the 8th International Conference on Language Resources and Evaluation, Istanbul, Turkey."},{"key":"ref_36","unstructured":"Barr\u00f3n-Cede\u00f1o, A., M\u00e0rquez-Villodre, L., Henr\u00edquez-Quintana, C.A., Formiga-Fanals, L., Romero-Merino, E., and May, J. (2013, January 3\u20139). Identifying useful human correction feedback from an on-line machine translation service. Proceedings of the 23rd International Joint Conference on Artificial Intelligence, Beijing, China."},{"key":"ref_37","unstructured":"Kordoni, V., Birch, L., Buliga, I., Cholakov, K., Egg, M., Gaspari, F., Georgakopoulou, Y., Gialama, M., Hendrickx, I.H.E., and Jermol, M. (June, January 30). TraMOOC (Translation for Massive Open Online Courses): Providing Reliable MT for MOOCs. Proceedings of the 19th annual conference of the European Association for Machine Translation (EAMT), Riga, Latvia."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Sennrich, R., Firat, O., Cho, K., Birch-Mayne, A., Haddow, B., Hitschler, J., Junczys-Dowmunt, M., L\u00e4ubli, S., Miceli Barone, A., and Mokry, J. (2017, January 3\u20137). Nematus: A toolkit for neural machine translation. Proceedings of the EACL 2017 Software Demonstrations, Valencia, Spain.","DOI":"10.18653\/v1\/E17-3017"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Miceli-Barone, A.V., Haddow, B., Germann, U., and Sennrich, R. (2017, January 9\u201311). Regularization techniques for ne-tuning in neural machine translation. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark.","DOI":"10.18653\/v1\/D17-1156"},{"key":"ref_40","first-page":"203","article-title":"Comparative evaluation of string similarity measures for automatic language classification","volume":"Volume 69","author":"Mikros","year":"2015","journal-title":"Sequences in Language and Text"},{"key":"ref_41","unstructured":"Broder, A.Z. (1997, January 11\u201313). On the resemblance and containment of documents. Proceedings of the Compression and Complexity of Sequences 1997, Washington, DC, USA."},{"key":"ref_42","unstructured":"Pouliquen, B., Steinberger, R., and Ignat, C. (2003, January 10\u201313). Automatic identification of document translations in large multilingual document collections. Proceedings of the International Conference Recent Advances in Natural Language Processing (RANLP), Borovets, Bulgaria."},{"key":"ref_43","unstructured":"(2018, October 08). Deep Learning for Java. Available online: https:\/\/deeplearning4j.org\/."},{"key":"ref_44","first-page":"250","article-title":"A study on WEKA tool for data preprocessing, classification and clustering","volume":"2","author":"Singhal","year":"2013","journal-title":"IJITEE"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1080\/08839510500313653","article-title":"Evaluation of classifiers for an uneven class distribution problem","volume":"20","author":"Daskalaki","year":"2006","journal-title":"Appl. Artif. Intell."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1613\/jair.953","article-title":"SMOTE: Synthetic minority over-sampling technique","volume":"16","author":"Chawla","year":"2002","journal-title":"J. Artif. Intell. Res."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Kuhn, M., and Kjell, J. (2013). Applied Predictive Modeling, Springer.","DOI":"10.1007\/978-1-4614-6849-3"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Zhang, D., and Lee, W.S. (2006, January 20\u201323). Extracting key-substring-group features for text classification. Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Philadelphia, PA, USA.","DOI":"10.1145\/1150402.1150455"},{"key":"ref_49","unstructured":"\u0160ili\u0107, A., Chauchat, J.H., Ba\u0161i\u0107, B.D., and Morin, A. (2007, January 3\u20137). N-grams and morphological normalization in text classification: A comparison on a croatian-english parallel corpus. Proceedings of the Portuguese Conference on Artificial Intelligence, Guimar\u00e3es, Portugal."},{"key":"ref_50","unstructured":"Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., and Isard, M. (2016, January 2\u20134). Tensorflow: A system for large-scale machine learning. Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation, Savannah, GA, USA."},{"key":"ref_51","unstructured":"Kovalev, V., Kalinovsky, A., and Kovalev, S. (2016). Deep Learning with Theano, Torch, Caffe, Tensorflow, and Deeplearning4j: Which One Is the Best in Speed and Accuracy?, Springer."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/12\/1\/26\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:27:14Z","timestamp":1760185634000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/12\/1\/26"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,1,18]]},"references-count":51,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2019,1]]}},"alternative-id":["a12010026"],"URL":"https:\/\/doi.org\/10.3390\/a12010026","relation":{},"ISSN":["1999-4893"],"issn-type":[{"type":"electronic","value":"1999-4893"}],"subject":[],"published":{"date-parts":[[2019,1,18]]}}}