{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T17:54:03Z","timestamp":1783014843069,"version":"3.54.6"},"reference-count":34,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,3,31]],"date-time":"2023-03-31T00:00:00Z","timestamp":1680220800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,3,31]]},"abstract":"<jats:p>Deep Learning is one of the most promising technologies compared to other methods in the context of machine translation. It has been proven to achieve impressive results on large amounts of parallel data for well-endowed languages. Nevertheless, for low-resource languages such as the Arabic Dialects, Deep Learning models failed due to the lack of available parallel corpora. In this article, we present a method to create a parallel corpus to build an effective NMT model able to translate into MSA, Tunisian Dialect texts present in social networks. For this, we propose a set of data augmentation methods aiming to increase the size of the state-of-the-art parallel corpus. By evaluating the impact of this step, we noticed that it has effectively boosted both the size and the quality of the corpus. Then, using the resulted corpus, we compare the effectiveness of CNN, RNN and transformers models to translate Tunisian Dialect into MSA. Experiments show that a better translation is achieved by the transformer model with a BLEU score of 60 vs., respectively, 33.36 and 53.98 with RNN and CNN models.<\/jats:p>","DOI":"10.1145\/3568674","type":"journal-article","created":{"date-parts":[[2022,11,2]],"date-time":"2022-11-02T13:02:51Z","timestamp":1667394171000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["Hybrid Pipeline for Building Arabic Tunisian Dialect-standard Arabic Neural Machine Translation Model from Scratch"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4674-7671","authenticated-orcid":false,"given":"Sam\u00e9h","family":"Kchaou","sequence":"first","affiliation":[{"name":"University of Sfax, Tunisia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2314-6064","authenticated-orcid":false,"given":"Rahma","family":"Boujelbane","sequence":"additional","affiliation":[{"name":"University of Sfax, Tunisia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4868-657X","authenticated-orcid":false,"given":"Lamia","family":"Hadrich","sequence":"additional","affiliation":[{"name":"University of Sfax, Tunisia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,4,14]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICICS49469.2020.239505"},{"key":"e_1_3_2_3_2","first-page":"167","volume-title":"Mach. Translat. 32 (2018),","author":"Karakanta Alina","year":"2018","unstructured":"Alina Karakanta, Jon Dehdari, and Josef van Genabith. 2018. Neural machine translation for low-resource languages without parallel corpora. Mach. Translat. 32 (2018),167\u2013189."},{"key":"e_1_3_2_4_2","first-page":"347","volume-title":"Machine Learning and Data Mining in Pattern Recognition","author":"Al-Ani Ebtesam H. Almansor and Ahmed","year":"2018","unstructured":"Ebtesam H. Almansor and Ahmed Al-Ani. 2018. A hybrid neural machine translation technique for translating low resource languages. In Machine Learning and Data Mining in Pattern Recognition. Springer International Publishing, Cham, 347\u2013356."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.26615\/978-954-452-049-6_008"},{"key":"e_1_3_2_6_2","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2016. Neural Machine Translation by Jointly Learning to Align and Translate. arxiv:1409.0473 [cs.CL]."},{"key":"e_1_3_2_7_2","first-page":"Dec. 10 (2018)","volume-title":"Computat. Intell. Neurosci.","author":"Baniata Laith H.","year":"2018","unstructured":"Laith H. Baniata, Seyoung Park, and Seong-Bae Park. 2018. A neural machine translation model for Arabic dialects that utilizes multitask learning (MTL). Computat. Intell. Neurosci.Dec. 10 (2018)."},{"key":"e_1_3_2_8_2","first-page":"1240","volume-title":"Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC\u201914)","author":"Bouamor Houda","year":"2014","unstructured":"Houda Bouamor, Nizar Habash, and Kemal Oflazer. 2014. A multidialectal parallel corpus of Arabic. In Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC\u201914). European Language Resources Association (ELRA), 1240\u20131245. Retrieved from http:\/\/www.lrec-conf.org\/proceedings\/lrec2014\/pdf\/523_Paper.pdf."},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the 11th Language Resources and Evaluation Conference","author":"Bouamor Houda","year":"2018","unstructured":"Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, and Kemal Oflazer. 2018. The MADAR Arabic dialect corpus and lexicon. In Proceedings of the 11th Language Resources and Evaluation Conference. European Language Resource Association. Retrieved from https:\/\/www.aclweb.org\/anthology\/L18-1535."},{"key":"e_1_3_2_10_2","first-page":"419","volume-title":"Proceedings of the 6th International Joint Conference on Natural Language Processing","author":"Boujelbane Rahma","year":"2013","unstructured":"Rahma Boujelbane, Mariem Ellouze Khemekhem, and Lamia Hadrich Belguith. 2013. Mapping rules for building a Tunisian dialect lexicon and generating corpora. In Proceedings of the 6th International Joint Conference on Natural Language Processing. Asian Federation of Natural Language Processing, 419\u2013428. Retrieved from https:\/\/www.aclweb.org\/anthology\/I13-1048."},{"key":"e_1_3_2_11_2","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Chen Kehai","year":"2017","unstructured":"Kehai Chen, Rui Wang, Masao Utiyama, Lemao Liu, Akihiro Tamura, Eiichiro Sumita, and Tiejun Zhao. 2017. Neural machine translation with source dependency representation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics."},{"key":"e_1_3_2_12_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arxiv:1810.04805 [cs.CL]."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/MoWNet.2016.7496606"},{"key":"e_1_3_2_14_2","volume-title":"Proceedings of the Association for Computational Linguistics","author":"Fei Gao","year":"2019","unstructured":"Gao Fei, Zhu Jinhua, Wu Lijun, Xia Yingce, Qin Tao, Cheng Xueqi, Zhou Wengang, and Liu Tie-Yan. 2019. Soft contextual data augmentation for neural machine translation. In Proceedings of the Association for Computational Linguistics."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2015.2464687"},{"key":"e_1_3_2_16_2","volume-title":"Proceedings of the MT Summit","author":"Hamdi Ahmed","year":"2013","unstructured":"Ahmed Hamdi, Rahma Boujelbane, Nizar Habash, and Alexis Nasr. 2013. The effects of factorizing root and pattern mapping in bidirectional Tunisian\u2013standard Arabic machine translation. In Proceedings of the MT Summit. Retrieved from https:\/\/hal.archives-ouvertes.fr\/hal-00908761."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-3627"},{"key":"e_1_3_2_18_2","article-title":"Corpus augmentation by sentence segmentation for low-resource neural machine translation","volume":"1905","author":"Jinyi Zhang","year":"2019","unstructured":"Zhang Jinyi and Matsumoto Tadahiro. 2019. Corpus augmentation by sentence segmentation for low-resource neural machine translation. CoRR abs\/1905.08945 (2019).","journal-title":"CoRR"},{"key":"e_1_3_2_19_2","first-page":"26","volume-title":"Proceedings of the 29th Pacific Asia Conference on Language, Information and Computation.","author":"Meftouh Karima","year":"2015","unstructured":"Karima Meftouh, Salima Harrat, S. Jamoussi, M. Abbas, and Kamel Sma\u00efli. 2015. Machine translation experiments on PADIC: A parallel Arabic dialect corpus. In Proceedings of the 29th Pacific Asia Conference on Language, Information and Computation.26\u201334."},{"key":"e_1_3_2_20_2","first-page":"200","volume-title":"Proceedings of the 5th Arabic Natural Language Processing Workshop","author":"Kchaou Sam\u00e9h","year":"2020","unstructured":"Sam\u00e9h Kchaou, Rahma Boujelbane, and Lamia Hadrich-Belguith. 2020. Parallel resources for Tunisian Arabic dialect translation. In Proceedings of the 5th Arabic Natural Language Processing Workshop. Association for Computational Linguistics, 200\u2013206. Retrieved from https:\/\/www.aclweb.org\/anthology\/2020.wanlp-1.18."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-4012"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Julia Kreutzer Jasmijn Bastings and Stefan Riezler. 2020. Joey NMT: A Minimalist NMT Toolkit for Novices. arxiv:1907.12484 [cs.CL].","DOI":"10.18653\/v1\/D19-3019"},{"key":"e_1_3_2_23_2","unstructured":"Surafel M. Lakew Mauro Cettolo and Marcello Federico. 2018. A Comparison of Transformer and Recurrent Neural Networks on Multilingual Neural Machine Translation. arxiv:1806.06957 [cs.CL]."},{"key":"e_1_3_2_24_2","first-page":"2078","volume-title":"Information11, 255 (2020),","author":"Li Y.","year":"2020","unstructured":"Y. Li, X. Li, Y. Yang, and R. Dong. 2020. A diverse data augmentation strategy for low-resource neural machine translation. Information11, 255 (2020),2078\u20132489."},{"key":"e_1_3_2_25_2","first-page":"567","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics.","author":"Marzieh Fadaee","year":"2017","unstructured":"Fadaee Marzieh, Bisazza Arianna, and Monz Christof. 2017. Data augmentation for low-resource neural machine translation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics.567\u2013573."},{"key":"e_1_3_2_26_2","unstructured":"Adam Paszke Sam Gross Francisco Massa Adam Lerer James Bradbury Gregory Chanan Trevor Killeen Zeming Lin Natalia Gimelshein Luca Antiga Alban Desmaison Andreas K\u00f6pf Edward Yang Zach DeVito Martin Raison Alykhan Tejani Sasank Chilamkurthy Benoit Steiner Lu Fang Junjie Bai and Soumith Chintala. 2019. PyTorch: An Imperative Style High-Performance Deep Learning Library. arxiv:1912.01703 [cs.LG]."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.winlp-1.40"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-2509"},{"key":"e_1_3_2_29_2","first-page":"385","volume-title":"Proceedings of COLING 2012: Demonstration Papers","author":"Salloum Wael","year":"2012","unstructured":"Wael Salloum and Nizar Habash. 2012. Elissa: A dialectal to standard Arabic machine translation system. In Proceedings of COLING 2012: Demonstration Papers. The COLING 2012 Organizing Committee, 385\u2013392. Retrieved from https:\/\/www.aclweb.org\/anthology\/C12-3048."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-2074"},{"key":"e_1_3_2_31_2","unstructured":"Bashar Talafha Mohammad Ali Muhy Eddin Za\u2019ter Haitham Seelawi Ibraheem Tuffaha Mostafa Samir Wael Farhan and Hussein T. Al-Natsheh. 2020. Multi-dialect Arabic BERT for Country-level Dialect Identification. arxiv:2007.05612 [cs.CL]."},{"key":"e_1_3_2_32_2","volume-title":"Proceedings of the 3rd Workshop on Technologies for MT of Low Resource Languages","author":"Tapo Allahsera Auguste","year":"2020","unstructured":"Allahsera Auguste Tapo, Bakary Coulibaly, S\u00e9bastien Diarra, Christopher Homan, Julia Kreutzer, Sarah Luger, Arthur Nagashima, Marcos Zampieri, and Michael Leventhal. 2020. Neural machine translation for extremely low-resource African languages: A case study on Bambara. In Proceedings of the 3rd Workshop on Technologies for MT of Low Resource Languages. Association for Computational Linguistics."},{"key":"e_1_3_2_33_2","article-title":"Attention is all you need","volume":"1706","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. CoRR abs\/1706.03762 (2017).","journal-title":"CoRR"},{"key":"e_1_3_2_34_2","unstructured":"Shijie Wu and Mark Dredze. 2019. Beto Bentz Becas: The Surprising Cross-lingual Effectiveness of BERT. arxiv:1904.09077 [cs.CL]."},{"key":"e_1_3_2_35_2","first-page":"147","article-title":"Morphological disambiguation of Tunisian dialect","volume":"29","author":"Zribi In\u00e8s","year":"2017","unstructured":"In\u00e8s Zribi, M. Ellouze, L. Belguith, and P. Blache. 2017. Morphological disambiguation of Tunisian dialect. J. King Saud Univ. Comput. Inf. Sci. 29 (2017), 147\u2013155.","journal-title":"J. King Saud Univ. Comput. Inf. Sci."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3568674","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3568674","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:08:34Z","timestamp":1750183714000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3568674"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,31]]},"references-count":34,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,3,31]]}},"alternative-id":["10.1145\/3568674"],"URL":"https:\/\/doi.org\/10.1145\/3568674","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,31]]},"assertion":[{"value":"2022-01-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-10-05","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}