{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T14:03:42Z","timestamp":1779199422276,"version":"3.51.4"},"reference-count":38,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2009,12,1]],"date-time":"2009-12-01T00:00:00Z","timestamp":1259625600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000185","name":"Defense Advanced Research Projects Agency","doi-asserted-by":"publisher","award":["HR0011-06-C-0022"],"award-info":[{"award-number":["HR0011-06-C-0022"]}],"id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Transactions on Asian Language Information Processing"],"published-print":{"date-parts":[[2009,12]]},"abstract":"<jats:p>The Arabic language presents a number of challenges for speech recognition, arising in part from the significant differences in the spoken and written forms, in particular the conventional form of texts being non-vowelized. Being a highly inflected language, the Arabic language has a very large lexical variety and typically with several possible (generally semantically linked) vowelizations for each written form. This article summarizes research carried out over the last few years on speech-to-text transcription of broadcast data in Arabic. The initial research was oriented toward processing of broadcast news data in Modern Standard Arabic, and has since been extended to address a larger variety of broadcast data, which as a consequence results in the need to also be able to handle dialectal speech. While standard techniques in speech recognition have been shown to apply well to the Arabic language, taking into account language specificities help to significantly improve system performance.<\/jats:p>","DOI":"10.1145\/1644879.1644885","type":"journal-article","created":{"date-parts":[[2010,1,5]],"date-time":"2010-01-05T15:05:08Z","timestamp":1262703908000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Automatic Speech-to-Text Transcription in Arabic"],"prefix":"10.1145","volume":"8","author":[{"given":"Lori","family":"Lamel","sequence":"first","affiliation":[{"name":"LIMSI-CNRS"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Abdelkhalek","family":"Messaoudi","sequence":"additional","affiliation":[{"name":"LIMSI-CNRS"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jean-Luc","family":"Gauvain","sequence":"additional","affiliation":[{"name":"LIMSI-CNRS"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2009,12]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the European Conference on Speech Technology (EuroSpeech\u201997)","author":"Adda G.","unstructured":"Adda , G. , Adda-Decker , M. , Gauvain , J.-L. , and Lamel , L . 1997. Text normalization and speech recognition in French . In Proceedings of the European Conference on Speech Technology (EuroSpeech\u201997) . 2711--2714. Adda, G., Adda-Decker, M., Gauvain, J.-L., and Lamel, L. 1997. Text normalization and speech recognition in French. In Proceedings of the European Conference on Speech Technology (EuroSpeech\u201997). 2711--2714."},{"key":"e_1_2_1_2_1","volume-title":"Proceedings of the European Conference on Speech Technology (EuroSpeech\u201903)","author":"Adda-Decker M.","year":"2003","unstructured":"Adda-Decker , M. 2003 . A corpus-based decompounding algorithm for German lexical modeling in LVCSR . In Proceedings of the European Conference on Speech Technology (EuroSpeech\u201903) . 257--260. Adda-Decker, M. 2003. A corpus-based decompounding algorithm for German lexical modeling in LVCSR. In Proceedings of the European Conference on Speech Technology (EuroSpeech\u201903). 257--260."},{"key":"e_1_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Adda-Decker M. and Lamel L. 2000. The use of lexica in automatic speech recognition. F. Van Eynde and D. Gibbon Eds. Kluwer Academic Publishers 235--266.  Adda-Decker M. and Lamel L. 2000. The use of lexica in automatic speech recognition . F. Van Eynde and D. Gibbon Eds. Kluwer Academic Publishers 235--266.","DOI":"10.1007\/978-94-010-9458-0_8"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the 9th European Conference on Speech Communication and Technology (EuroSpeech\u201905)","author":"Afify M.","unstructured":"Afify , M. , Nguyen , L. , Xiang , B. , Abdou , S. , and Makhoul , J . 2005. Recent progress in Arabic broadcast news transcription at BBN . In Proceedings of the 9th European Conference on Speech Communication and Technology (EuroSpeech\u201905) . 1637--1640. Afify, M., Nguyen, L., Xiang, B., Abdou, S., and Makhoul, J. 2005. Recent progress in Arabic broadcast news transcription at BBN. In Proceedings of the 9th European Conference on Speech Communication and Technology (EuroSpeech\u201905). 1637--1640."},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201996)","author":"Anastasakos T.","unstructured":"Anastasakos , T. , McDonough , J. , Schwartz , R. , and Makhoul , J . 1996. A compact model for speaker adaptation training . In Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201996) . 1137--1140. Anastasakos, T., McDonough, J., Schwartz, R., and Makhoul, J. 1996. A compact model for speaker adaptation training. In Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201996). 1137--1140."},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of ICASSP.","volume":"1","author":"Billa J.","unstructured":"Billa , J. , Noamany , N. , Srivastava , A. , Liu , D. , Stone , R. , Xu , J. , Makhoul , J. , and Kubala , F . 2002. Audio indexing of Arabic broadcast news . In Proceedings of ICASSP. Vol. 1 . 5--8. Billa, J., Noamany, N., Srivastava, A., Liu, D., Stone, R., Xu, J., Makhoul, J., and Kubala, F. 2002. Audio indexing of Arabic broadcast news. In Proceedings of ICASSP. Vol. 1. 5--8."},{"key":"e_1_2_1_7_1","unstructured":"Buckwalter T. 2004. Arabic morphology analysis. www.qamus.org\/morphology.htm.  Buckwalter T. 2004. Arabic morphology analysis. www.qamus.org\/morphology.htm."},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201900)","author":"Carki K.","unstructured":"Carki , K. , Geutner , P. , and Schultz , T . 2000. Turkish LVCSR: Towards better speech recognition for agglutinative languages . In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201900) . 3688--3691. Carki, K., Geutner, P., and Schultz, T. 2000. Turkish LVCSR: Towards better speech recognition for agglutinative languages. In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201900). 3688--3691."},{"key":"e_1_2_1_9_1","volume-title":"Eds","author":"Chou W.","year":"2003","unstructured":"Chou , W. and Juang , F. , Eds . 2003 . Pattern Recognition in Speech and Language Processing. CRC Press . Chou, W. and Juang, F., Eds. 2003. Pattern Recognition in Speech and Language Processing. CRC Press."},{"key":"e_1_2_1_10_1","unstructured":"Creutz M. and Lagus K. 2005. Unsupervised morpheme segmentation and morphology induction from text corpora using Morfessor 1.0. Tech. rep. Computer and Information Science Tech. rep. A81. Helsinki University of Technology.  Creutz M. and Lagus K. 2005. Unsupervised morpheme segmentation and morphology induction from text corpora using Morfessor 1.0. Tech. rep. Computer and Information Science Tech. rep. A81. Helsinki University of Technology."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU.1997.659110"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-87391-4_39"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the Annual Conference of the International Speech Communication Association (InterSpeech\u201908)","author":"Fousek P.","unstructured":"Fousek , P. , Lamel , L. , and Gauvain , J . -L. 2008b. Transcribing broadcast data using MLP features . In Proceedings of the Annual Conference of the International Speech Communication Association (InterSpeech\u201908) . 1433--1436. Fousek, P., Lamel, L., and Gauvain, J.-L. 2008b. Transcribing broadcast data using MLP features. In Proceedings of the Annual Conference of the International Speech Communication Association (InterSpeech\u201908). 1433--1436."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.880079"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201998)","author":"Gauvain J.-L.","unstructured":"Gauvain , J.-L. , Lamel , L. , and Adda , G . 1998. Partitioning and transcription of broadcast news data . In Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201998) . 1335--1338. Gauvain, J.-L., Lamel, L., and Adda, G. 1998. Partitioning and transcription of broadcast news data. In Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201998). 1335--1338."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-6393(01)00061-9"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 9th European Conference on Speech Communication and Technology (InterSpeech\u201905)","author":"Ghaoui A.","unstructured":"Ghaoui , A. , Yvon , F. , Mokbel , C. , and Chollet , G . 2005. On the use of morphological constraints in n-gram statistical language model . In Proceedings of the 9th European Conference on Speech Communication and Technology (InterSpeech\u201905) . 1281--1284. Ghaoui, A., Yvon, F., Mokbel, C., and Chollet, G. 2005. On the use of morphological constraints in n-gram statistical language model. In Proceedings of the 9th European Conference on Speech Communication and Technology (InterSpeech\u201905). 1281--1284."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1162\/089120101750300490"},{"key":"e_1_2_1_20_1","volume-title":"Arabic Gigaword 3rd Ed. LDC LDC2007T40","author":"Graff D.","year":"2007","unstructured":"Graff , D. 2007 . Arabic Gigaword 3rd Ed. LDC LDC2007T40 . Graff, D. 2007. Arabic Gigaword 3rd Ed. LDC LDC2007T40."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201908)","author":"Gr\u00e9zl F.","unstructured":"Gr\u00e9zl , F. and Fousek , P . 2008. Optimizing bottle-neck features for LVCSR . In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201908) . 4729--4732. Gr\u00e9zl, F. and Fousek, P. 2008. Optimizing bottle-neck features for LVCSR. In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201908). 4729--4732."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.2307\/411036"},{"key":"e_1_2_1_23_1","volume-title":"et al","author":"Kirchhoff K.","year":"2002","unstructured":"Kirchhoff , K. et al . 2002 . Novel approaches to Arabic speech recognition. Tech. rep. (final report from the 2002 JHU summer workshop), John-Hopkins University , Baltimore, MD. Kirchhoff, K. et al. 2002. Novel approaches to Arabic speech recognition. Tech. rep. (final report from the 2002 JHU summer workshop), John-Hopkins University, Baltimore, MD."},{"key":"e_1_2_1_24_1","volume-title":"-L","author":"Lamel L.","year":"2003","unstructured":"Lamel , L. and Gauvain , J . -L . 2003 . Speech recognition. In OUP Handbook on Computational Linguistics, R. Mitkov, Ed. Oxford University Press , 305--322. Lamel, L. and Gauvain, J.-L. 2003. Speech recognition. In OUP Handbook on Computational Linguistics, R. Mitkov, Ed. Oxford University Press, 305--322."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-85287-2_2"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1006\/csla.1995.0010"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the European Conference on Speech Technology (EuroSpeech\u201999)","author":"Mangu L.","unstructured":"Mangu , L. , Brill , E. , and Stolcke , A . 1999. Finding consensus among words: Lattice-based word error minimization . In Proceedings of the European Conference on Speech Technology (EuroSpeech\u201999) . 495--498. Mangu, L., Brill, E., and Stolcke, A. 1999. Finding consensus among words: Lattice-based word error minimization. In Proceedings of the European Conference on Speech Technology (EuroSpeech\u201999). 495--498."},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201906)","author":"Messaoudi A.","unstructured":"Messaoudi , A. , Gauvain , J.-L. , and Lamel , L . 2006. Arabic broadcast news transcription using a one million word vocalized vocabulary . In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201906) . 1093--1096. Messaoudi, A., Gauvain, J.-L., and Lamel, L. 2006. Arabic broadcast news transcription using a one million word vocalized vocabulary. In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201906). 1093--1096."},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201904)","author":"Messaoudi A.","unstructured":"Messaoudi , A. , Lamel , L. , and Gauvain , J . -L. 2004. Transcription of Arabic broadcast news . In Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201904) . 521--524. Messaoudi, A., Lamel, L., and Gauvain, J.-L. 2004. Transcription of Arabic broadcast news. In Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201904). 521--524."},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the 9th European Conference on Speech Communication and Technology (InterSpeech\u201905)","author":"Messaoudi A.","unstructured":"Messaoudi , A. , Lamel , L. , and Gauvain , J . -L. 2005. Modeling vowels for Arabic BN transcription . In Proceedings of the 9th European Conference on Speech Communication and Technology (InterSpeech\u201905) . 1633--1636. Messaoudi, A., Lamel, L., and Gauvain, J.-L. 2005. Modeling vowels for Arabic BN transcription. In Proceedings of the 9th European Conference on Speech Communication and Technology (InterSpeech\u201905). 1633--1636."},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the Conference on New Methods in Language Processing (NeMLaP\u201994)","author":"Schmid H.","year":"1994","unstructured":"Schmid , H. 1994 . Probabilistic part-of-speech tagging using decision trees . In Proceedings of the Conference on New Methods in Language Processing (NeMLaP\u201994) . Schmid, H. 1994. Probabilistic part-of-speech tagging using decision trees. In Proceedings of the Conference on New Methods in Language Processing (NeMLaP\u201994)."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2006.09.003"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.3115\/1220575.1220601"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2006.879807"},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201904)","author":"Vergyri D.","unstructured":"Vergyri , D. , Kirchhoff , K. , Duh , K. , and Stolcke , A . 2004. Morphology-based language modeling for Arabic speech recognition . In Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201904) . 1252--1255. Vergyri, D., Kirchhoff, K., Duh, K., and Stolcke, A. 2004. Morphology-based language modeling for Arabic speech recognition. In Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201904). 1252--1255."},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201900)","author":"Whittaker E.","unstructured":"Whittaker , E. and Woodland , P . 2000. Particle-based language modeling . In Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201900) . Whittaker, E. and Woodland, P. 2000. Particle-based language modeling. In Proceedings of the International Conference on Speech and Language Processing (ICSLP\u201900)."},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201906)","author":"Xiang B.","unstructured":"Xiang , B. , Nguyen , K. , Nguyen , L. , Schwartz , R. , and Makhoul , J . 2006. Morphological decomposition for Arabic broadcast news transcription . In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201906) . 1089--1092. Xiang, B., Nguyen, K., Nguyen, L., Schwartz, R., and Makhoul, J. 2006. Morphological decomposition for Arabic broadcast news transcription. In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201906). 1089--1092."},{"key":"e_1_2_1_38_1","volume-title":"eds","author":"Young S.","year":"2000","unstructured":"Young , S. and Bloothooft , G. , eds . 2000 . Corpus-Based Methods in Language and Speech Processing. Kluwer Academic Publishers . Young, S. and Bloothooft, G., eds. 2000. Corpus-Based Methods in Language and Speech Processing. Kluwer Academic Publishers."},{"key":"e_1_2_1_39_1","volume-title":"Proceedings of the 9th European Conference on Speech Communication and Technology (InterSpeech\u201905)","author":"Zhu Q.","unstructured":"Zhu , Q. , Stolcke , A. , Chen , B.-Y. , and Morgan , N . 2005. Using MLP features in SRI\u2019s conversational speech recognition system . In Proceedings of the 9th European Conference on Speech Communication and Technology (InterSpeech\u201905) . 2141--2144. Zhu, Q., Stolcke, A., Chen, B.-Y., and Morgan, N. 2005. Using MLP features in SRI\u2019s conversational speech recognition system. In Proceedings of the 9th European Conference on Speech Communication and Technology (InterSpeech\u201905). 2141--2144."}],"container-title":["ACM Transactions on Asian Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1644879.1644885","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1644879.1644885","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T12:41:18Z","timestamp":1750250478000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1644879.1644885"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,12]]},"references-count":38,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2009,12]]}},"alternative-id":["10.1145\/1644879.1644885"],"URL":"https:\/\/doi.org\/10.1145\/1644879.1644885","relation":{},"ISSN":["1530-0226","1558-3430"],"issn-type":[{"value":"1530-0226","type":"print"},{"value":"1558-3430","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,12]]},"assertion":[{"value":"2009-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-12-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}