{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:06:06Z","timestamp":1750309566400,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":32,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,12,16]],"date-time":"2024-12-16T00:00:00Z","timestamp":1734307200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,12,16]]},"DOI":"10.1145\/3677389.3702509","type":"proceedings-article","created":{"date-parts":[[2025,3,13]],"date-time":"2025-03-13T16:53:52Z","timestamp":1741884832000},"page":"1-5","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Arabic Text Enhancement with GPT for Digital Libraries"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4833-8882","authenticated-orcid":false,"given":"Luca","family":"Sala","sequence":"first","affiliation":[{"name":"Department of Engineering Enzo Ferrari, University of Modena and Reggio Emilia, Modena, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5556-1827","authenticated-orcid":false,"given":"Giovanni","family":"Sullutrone","sequence":"additional","affiliation":[{"name":"University of Modena and Reggio Emilia, Modena, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8087-6587","authenticated-orcid":false,"given":"Sonia","family":"Bergamaschi","sequence":"additional","affiliation":[{"name":"University of Modena and Reggio Emilia, Modena, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4643-6128","authenticated-orcid":false,"given":"Riccardo Amerigo","family":"Vigliermo","sequence":"additional","affiliation":[{"name":"University of Modena and Reggio Emilia, Modena, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,3,13]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2024. GPT-4V(ision). https:\/\/openai.com\/index\/gpt-4v-system-card\/ Accessed: 2024-05-09."},{"key":"e_1_3_2_1_2_1","unstructured":"2024. Tesseract OCR. https:\/\/tesseract-ocr.github.io\/Accessed: 2024-05-09."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.11591\/ijeecs.v26.i2.pp754-763"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008147606320"},{"key":"e_1_3_2_1_5_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_3_2_1_6_1","first-page":"175","article-title":"OCR Error Correction Using Statistical Machine","volume":"7","author":"Afli Haithem","year":"2016","unstructured":"Haithem Afli, Lo\u00efc Barrault, and Holger Schwenk. 2016. OCR Error Correction Using Statistical Machine Translation. Int. J. Comput. Linguistics Appl. 7, 1 (2016), 175--191.","journal-title":"Translation. Int. J. Comput. Linguistics Appl."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CSIT.2016.7549460"},{"key":"e_1_3_2_1_8_1","unstructured":"Giorgio Banti. 1994. Introduzione alle lingue semitiche (Studi sul Vicino Oriente antico 2)."},{"key":"e_1_3_2_1_9_1","volume-title":"Ocr post-processing error correction algorithm using google online spelling suggestion. arXiv preprint arXiv:1204.0191","author":"Bassil Youssef","year":"2012","unstructured":"Youssef Bassil and Mohammad Alwani. 2012. Ocr post-processing error correction algorithm using google online spelling suggestion. arXiv preprint arXiv:1204.0191 (2012)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.3390\/s22113995"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i1.19885"},{"key":"e_1_3_2_1_12_1","first-page":"133","article-title":"Post-correction of Historical Text Transcripts with Large Language Models: An Exploratory Study","volume":"2024","author":"Boros Emanuela","year":"2024","unstructured":"Emanuela Boros, Maud Ehrmann, Matteo Romanello, Sven Najem-Meyer, and Fr\u00e9d\u00e9ric Kaplan. 2024. Post-correction of Historical Text Transcripts with Large Language Models: An Exploratory Study. LaTeCH-CLfL 2024 (2024), 133--159.","journal-title":"LaTeCH-CLfL"},{"key":"e_1_3_2_1_13_1","volume-title":"Language models are few-shot learners. arXiv preprint arXiv:2005.14165","author":"Brown Tom B","year":"2020","unstructured":"Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020)."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-43427-3_35"},{"key":"e_1_3_2_1_15_1","unstructured":"Olivier Durand Angela Langone Giuliano Mion et al. 2010. Corso di arabo contemporaneo. Lingua standard. Hoepli."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3322905.3322908"},{"key":"e_1_3_2_1_17_1","unstructured":"Mahdi Hajiali. 2023. OCR Post-processing Using Large Language Models. (2023)."},{"key":"e_1_3_2_1_18_1","first-page":"143","article-title":"Toward a New Typology of Al-\u1e6cib\u0101q","volume":"55","author":"Hassanein Hamada","year":"2022","unstructured":"Hamada Hassanein. 2022. Toward a New Typology of Al-\u1e6cib\u0101q\" Antonymy\" in Qur'anic Arabic. Al-\u02bfArabiyya: Journal of the American Association of Teachers of Arabic 55, 1 (2022), 143--174.","journal-title":"Antonymy\" in Qur'anic Arabic. Al-\u02bfArabiyya: Journal of the American Association of Teachers of Arabic"},{"key":"e_1_3_2_1_19_1","first-page":"3019","article-title":"A rule-based post-processing approach to improve Persian OCR performance","volume":"27","author":"Khosrobeigi Zohreh","year":"2020","unstructured":"Zohreh Khosrobeigi, Hadi Veisi, HR Ahmadi, and Hanieh Shabanian. 2020. A rule-based post-processing approach to improve Persian OCR performance. Scientia Iranica 27, 6 (2020), 3019--3033.","journal-title":"Scientia Iranica"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/DAS.2016.44"},{"key":"e_1_3_2_1_21_1","volume-title":"Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa.","author":"Kojima Takeshi","year":"2022","unstructured":"Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems 35 (2022), 22199--22213."},{"key":"e_1_3_2_1_22_1","volume-title":"Better zero-shot reasoning with role-play prompting. arXiv preprint arXiv:2308.07702","author":"Kong Aobo","year":"2023","unstructured":"Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, and Xiaohang Dong. 2023. Better zero-shot reasoning with role-play prompting. arXiv preprint arXiv:2308.07702 (2023)."},{"key":"e_1_3_2_1_23_1","volume-title":"On the hidden mystery of ocr in large multimodal models. arXiv preprint arXiv:2305.07895","author":"Liu Yuliang","year":"2023","unstructured":"Yuliang Liu, Zhang Li, Biao Yang, Chunyuan Li, Xucheng Yin, Cheng-lin Liu, Lianwen Jin, and Xiang Bai. 2023. On the hidden mystery of ocr in large multimodal models. arXiv preprint arXiv:2305.07895 (2023)."},{"key":"e_1_3_2_1_24_1","volume-title":"Proceedings of the 19th The Conference on Information and Research science Connecting to Digital and Library science, IRCDL 2023","author":"Martoglia Riccardo","year":"2023","unstructured":"Riccardo Martoglia, Sonia Bergamaschi, Federico Ruozzi, Matteo Vanzini, Luca Sala, Riccardo Amerigo Vigliermo, et al. 2023. Knowledge extraction, management and long-term preservation of non-Latin cultural heritages-Digital Maktaba project presentation. In Proceedings of the 19th The Conference on Information and Research science Connecting to Digital and Library science, IRCDL 2023, Bari, Italy, February 23--24, 2023."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3476887.3476888"},{"key":"e_1_3_2_1_26_1","volume-title":"Large Language Models and Information Retrieval. Available at SSRN 4636121","author":"Pakhale Kalyani","year":"2023","unstructured":"Kalyani Pakhale. 2023. Large Language Models and Information Retrieval. Available at SSRN 4636121 (2023)."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksuci.2022.04.021"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2009.155"},{"key":"e_1_3_2_1_29_1","unstructured":"Jingqun Tang Chunhui Lin Zhen Zhao Shu Wei Binghong Wu Qi Liu Hao Feng Yang Li Siqi Wang Lei Liao et al. 2024. TextSquare: Scaling up Text-Centric Visual Instruction Tuning. arXiv preprint arXiv:2404.12803 (2024)."},{"key":"e_1_3_2_1_30_1","volume-title":"The cognitive processes involved in learning to read in Arabic. Reading and writing 17","author":"Taouka Miriam","year":"2004","unstructured":"Miriam Taouka and Max Coltheart. 2004. The cognitive processes involved in learning to read in Arabic. Reading and writing 17 (2004), 27--57."},{"key":"e_1_3_2_1_31_1","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Yonghui Wu Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew M Dai Anja Hauth et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023)."},{"key":"e_1_3_2_1_32_1","volume-title":"Large language models for generative information extraction: A survey. arXiv preprint arXiv:2312.17617","author":"Xu Derong","year":"2023","unstructured":"Derong Xu, Wei Chen, Wenjun Peng, Chao Zhang, Tong Xu, Xiangyu Zhao, Xian Wu, Yefeng Zheng, and Enhong Chen. 2023. Large language models for generative information extraction: A survey. arXiv preprint arXiv:2312.17617 (2023)."}],"event":{"name":"JCDL '24: 24th ACM\/IEEE Joint Conference on Digital Libraries","sponsor":["SIGIR ACM Special Interest Group on Information Retrieval","SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","IEEE TCDL"],"location":"Hong Kong China","acronym":"JCDL '24"},"container-title":["Proceedings of the 24th ACM\/IEEE Joint Conference on Digital Libraries"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3677389.3702509","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3677389.3702509","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:19:07Z","timestamp":1750295947000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3677389.3702509"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,16]]},"references-count":32,"alternative-id":["10.1145\/3677389.3702509","10.1145\/3677389"],"URL":"https:\/\/doi.org\/10.1145\/3677389.3702509","relation":{},"subject":[],"published":{"date-parts":[[2024,12,16]]},"assertion":[{"value":"2025-03-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}