{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,3]],"date-time":"2026-03-03T16:41:28Z","timestamp":1772556088800,"version":"3.50.1"},"reference-count":30,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,3,25]],"date-time":"2023-03-25T00:00:00Z","timestamp":1679702400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100008982","name":"Qatar National Research Fund","doi-asserted-by":"crossref","award":["NPRP13S-0112-200037"],"award-info":[{"award-number":["NPRP13S-0112-200037"]}],"id":[{"id":"10.13039\/100008982","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,4,30]]},"abstract":"<jats:p>With the advances in Natural Language Processing (NLP), the industry has been moving towards human-directed artificial intelligence (AI) solutions. Recently, chatbots and automated news generation have captured a lot of attention. The goal is to automatically generate readable text from tabular data or web data commonly represented in Resource Description Framework (RDF) format. The problem can then be formulated as Data-to-text (D2T) generation from structured non-linguistic data into human-readable natural language. Despite the significant work done for the English language, no efforts are being directed towards low-resource languages like the Arabic language. This work promotes the development of the first RDF data-to-text (D2T) generation system for the Arabic language while trying to address the low-resource limitation. We develop several models for the Arabic D2T task using transfer learning from large language models (LLM) such as AraBERT, AraGPT2, and mT5. These models include a baseline Bi-LSTM Sequence-to-Sequence (Seq2Seq) model, as well as encoder-decoder transformers like BERT2BERT, BERT2GPT, and T5. We then provide a detailed comparative study highlighting the strengths and limitations of these methods setting the stage for further advancement in the field. We also introduce a new Arabic dataset (AraWebNLG) that can be used for new model development in the field. To ensure a comprehensive evaluation, general-purpose automated metrics (BLEU and Perplexity scores) are used as well as task-specific human evaluation metrics related to the accuracy of the content selection and fluency of the generated text. The results highlight the importance of pre-training on a large corpus of Arabic data and show that transfer learning from AraBERT gives the best performance. Text-to-text pre-training using mT5 achieves second best performance results even with multilingual weights.<\/jats:p>","DOI":"10.1145\/3582262","type":"journal-article","created":{"date-parts":[[2023,1,25]],"date-time":"2023-01-25T11:54:25Z","timestamp":1674647665000},"page":"1-13","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Automated Generation of Human-readable Natural Arabic Text from RDF Data"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0033-4989","authenticated-orcid":false,"given":"Roudy","family":"Touma","sequence":"first","affiliation":[{"name":"American University of Beirut, Beirut, Lebanon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9954-7924","authenticated-orcid":false,"given":"Hazem","family":"Hajj","sequence":"additional","affiliation":[{"name":"American University of Beirut, Beirut, Lebanon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5206-2954","authenticated-orcid":false,"given":"Wassim","family":"El-Hajj","sequence":"additional","affiliation":[{"name":"American University of Beirut, Beirut, Lebanon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5688-7515","authenticated-orcid":false,"given":"Khaled","family":"Shaban","sequence":"additional","affiliation":[{"name":"Qatar University, Doha, Qatar"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,3,25]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Wissam Antoun Fady Baly and Hazem Hajj. 2020. AraBERT: Transformer-based model for Arabic language understanding. In LREC 2020 Workshop Language Resources and Evaluation Conference 11.16 May 2020 . 9."},{"key":"e_1_3_2_3_2","unstructured":"Wissam Antoun Fady Baly and Hazem Hajj. 2021. AraGPT2: Pre-Trained transformer for Arabic language generation. In Proceedings of the 6th Arabic Natural Language Processing Workshop . Association for Computational Linguistics Kyiv (Virtual) 196\u2013207."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2009.102"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1547"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"e_1_3_2_7_2","first-page":"1070","volume-title":"Proceedings of the 10th International Conference on Language Resources and Evaluation","author":"Darwish Kareem","year":"2016","unstructured":"Kareem Darwish and Hamdy Mubarak. 2016. Farasa: A new fast and accurate Arabic word segmenter. In Proceedings of the 10th International Conference on Language Resources and Evaluation. 1070\u20131074."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-4608"},{"key":"e_1_3_2_10_2","volume-title":"Proceedings of the 3rd International Workshop on Natural Language Generation from the Semantic Web","author":"Ferreira Thiago","year":"2020","unstructured":"Thiago Ferreira, Claire Gardent, Nikolai Ilinykh, Chris van der Lee, Simon Mille, Diego Moussallem, and Anastasia Shimorina. 2020. The 2020 bilingual, bi-directional webnlg+ shared task overview and evaluation results (webnlg+ 2020). In Proceedings of the 3rd International Workshop on Natural Language Generation from the Semantic Web."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1052"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-3518"},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","unstructured":"Mihir Kale and Abhinav Rastogi. 2020. Text-to-text pre-training for data-to-text tasks. In Proceedings of the 13th International Conference on Natural Language Generation Association for Computational Linguistics Dublin 97\u2013102. https:\/\/aclanthology.org\/2020.inlg-1.14.","DOI":"10.18653\/v1\/2020.inlg-1.14"},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","unstructured":"Guillaume Klein Yoon Kim Yuntian Deng Jean Senellart and Alexander Rush. 2017. OpenNMT: Open-source toolkit for Neural Machine Translation. In Proceedings of ACL 2017 System Demonstrations . Association for Computational Linguistics Vancouver 67\u201372. https:\/\/aclanthology.org\/P17-4012.","DOI":"10.18653\/v1\/P17-4012"},{"key":"e_1_3_2_15_2","first-page":"117","volume-title":"Proceedings of the 3rd International Workshop on Natural Language Generation from the Semantic Web","author":"Li Xintong","year":"2020","unstructured":"Xintong Li, Aleksandre Maskharashvili, Symon Jory Stevens-Guille, and Michael White. 2020. Leveraging large pretrained models for WebNLG 2020. In Proceedings of the 3rd International Workshop on Natural Language Generation from the Semantic Web. 117\u2013124."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1236"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-54832-2_13"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.89"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33016908"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-6557"},{"key":"e_1_3_2_21_2","unstructured":"Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog 1 8 (2019) 9."},{"key":"e_1_3_2_22_2","unstructured":"Colin Raffel Noam Shazeer Adam Roberts Katherine Lee Sharan Narang Michael Matena Yanqi Zhou Wei Li and Peter J. Liu. 2022. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21 1 (Jun 2022) 67 pages."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-45439-5_5"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.nlp4convai-1.20"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00313"},{"key":"e_1_3_2_26_2","first-page":"5998","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/2629489"},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","first-page":"289","DOI":"10.18653\/v1\/D19-5633","volume-title":"Proceedings of the 3rd Workshop on Neural Generation and Translation","author":"Werlen Lesly Miculicich","year":"2019","unstructured":"Lesly Miculicich Werlen, Marc Marone, and Hany Hassan. 2019. Selecting, planning, and rewriting: A modular approach for data-to-document generation and translation. In Proceedings of the 3rd Workshop on Neural Generation and Translation. 289\u2013296."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1239"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.41"},{"key":"e_1_3_2_31_2","unstructured":"Rong Ye Wenxian Shi Hao Zhou Zhongyu Wei and Lei Li. 2020. Variational template machine for data-to-text generation. In International Conference on Learning Representations . https:\/\/openreview.net\/forum?id=HkejNgBtPB."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582262","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3582262","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:32Z","timestamp":1750182692000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582262"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,25]]},"references-count":30,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,4,30]]}},"alternative-id":["10.1145\/3582262"],"URL":"https:\/\/doi.org\/10.1145\/3582262","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,25]]},"assertion":[{"value":"2022-06-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-01-16","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}