{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,7]],"date-time":"2026-02-07T19:18:40Z","timestamp":1770491920728,"version":"3.49.0"},"reference-count":55,"publisher":"Cambridge University Press (CUP)","issue":"3","license":[{"start":{"date-parts":[[2020,3,16]],"date-time":"2020-03-16T00:00:00Z","timestamp":1584316800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2021,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Arabic sentiment analysis models have recently employed compositional paragraph or sentence embedding features to represent the informal Arabic dialectal content. These embeddings are mostly composed via ordered, syntax-aware composition functions and learned within deep neural network architectures. With the differences in the syntactic structure and words\u2019 order among the Arabic dialects, a sentiment analysis system developed for one dialect might not be efficient for the others. Here we present syntax-ignorant, sentiment-specific n-gram embeddings for sentiment analysis of several Arabic dialects. The novelty of the proposed model is illustrated through its features and architecture. In the proposed model, the sentiment is expressed by embeddings, composed via the unordered additive composition function and learned within a shallow neural architecture. To evaluate the generated embeddings, they were compared with the state-of-the art word\/paragraph embeddings. This involved investigating their efficiency, as expressive sentiment features, based on the visualisation maps constructed for our n-gram embeddings and word2vec\/doc2vec. In addition, using several Eastern\/Western Arabic datasets of single-dialect and multi-dialectal contents, the ability of our embeddings to recognise the sentiment was investigated against word\/paragraph embeddings-based models. This comparison was performed within both shallow and deep neural network architectures and with two unordered composition functions employed. The results revealed that the introduced syntax-ignorant embeddings could represent single and combinations of different dialects efficiently, as our shallow sentiment analysis model, trained with the proposed n-gram embeddings, could outperform the word2vec\/doc2vec models and rival deep neural architectures consuming, remarkably, less training time.<\/jats:p>","DOI":"10.1017\/s135132492000008x","type":"journal-article","created":{"date-parts":[[2020,3,16]],"date-time":"2020-03-16T10:21:19Z","timestamp":1584354079000},"page":"315-338","update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":4,"title":["Syntax-ignorant N-gram embeddings for dialectal Arabic sentiment analysis"],"prefix":"10.1017","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7608-2765","authenticated-orcid":false,"given":"Hala","family":"Mulki","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hatem","family":"Haddad","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mourad","family":"Gridach","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ismail","family":"Babao\u011flu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2020,3,16]]},"reference":[{"key":"S135132492000008X_ref12","author":"Brustad","year":"2000"},{"key":"S135132492000008X_ref28","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1181"},{"key":"S135132492000008X_ref45","unstructured":"Socher, R. , Bengio, Y. and Manning, C. (2013). Deep learning for NLP. In Tutorial at Association of Computational Logistics (ACL) and North American Chapter of the Association of Computational Linguistics (NAACL)."},{"key":"S135132492000008X_ref15","unstructured":"Courbariaux, M. , Hubara, I. , Soudry, D. , El-Yaniv, R. and Bengio, Y. (2016). Binarized neural networks: Training deep neural networks with weights and activations constrained to+ 1 or-1. arXiv preprint arXiv:1602.02830"},{"key":"S135132492000008X_ref26","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1162"},{"key":"S135132492000008X_ref40","unstructured":"Refaee, E. and Rieser, V. (2014) An Arabic twitter corpus for subjectivity and sentiment analysis. In LREC, pp. 2268\u20132273."},{"key":"S135132492000008X_ref29","unstructured":"Kingma, D.P. and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980."},{"key":"S135132492000008X_ref27","unstructured":"Karmani, N. (2017). Tunisian Arabic Customer\u2019s Reviews Processing and Analysis for an Internet Supervision System . PhD Dissertation, National Engineering School of Sfax, Tunisia."},{"key":"S135132492000008X_ref23","doi-asserted-by":"crossref","unstructured":"Gormley, M.R. , Yu, M. and Dredze, M. (2015). Improved relation extraction with feature-rich compositional embedding models. arXiv preprint arXiv:1505.02419.","DOI":"10.18653\/v1\/D15-1205"},{"key":"S135132492000008X_ref31","unstructured":"Le, Q. and Mikolov, T. (2014). Distributed representations of sentences and documents. In Proceedings of the 31st International Conference on Machine Learning (ICML-14), pp. 1188\u20131196."},{"key":"S135132492000008X_ref30","doi-asserted-by":"publisher","DOI":"10.1613\/jair.4272"},{"key":"S135132492000008X_ref9","first-page":"2654","article-title":"Do deep nets really need to be deep?","author":"Ba","year":"2014","journal-title":"Advances in Neural Information Processing Systems"},{"key":"S135132492000008X_ref55","article-title":"Deep nets don\u2019t learn via memorization","author":"Krueger","year":"2017","journal-title":"Workshop track- ICLR 2017"},{"key":"S135132492000008X_ref7","doi-asserted-by":"publisher","DOI":"10.1109\/BigData.2016.7841054"},{"key":"S135132492000008X_ref4","unstructured":"Al-Rfou, R. , Perozzi, B. and Skiena, S. (2013) Polyglot: Distributed word representations for multilingual NLP. In Proceedings of the Seventeenth Conference on Computational Natural Language Learning, pp. 183\u2013192."},{"key":"S135132492000008X_ref34","doi-asserted-by":"crossref","unstructured":"Mitchell, J. and Lapata, M. (2010). Recursive deep models for semantic compositionality over a sentiment treebank. In Composition in Distributional Models of Semantics, 1388\u20131429.","DOI":"10.1111\/j.1551-6709.2010.01106.x"},{"key":"S135132492000008X_ref10","unstructured":"Banea, C. , Mihalcea, R. and Wiebe, J. (2010). Multilingual subjectivity: Are more languages better? In Proceedings of the 23rd International Conference on Computational Linguistics, Association for Computational Linguistics (ACL), pp. 28\u201336."},{"key":"S135132492000008X_ref6","first-page":"25","article-title":"AROMA: A recursive deep learning model for opinion mining in Arabic as a low resource language","volume":"16","author":"Al-Sallab","year":"2017","journal-title":"ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP)"},{"key":"S135132492000008X_ref43","unstructured":"Sayadi, K. , Liwicki, M. , Ingold, R. and Bui, M. (2016). Tunisian dialect and modern standard Arabic dataset for sentiment analysis : Tunisian election context. In 2nd International Conference on Arabic Computational Linguistics (acling), pp. 120\u2013150."},{"key":"S135132492000008X_ref22","unstructured":"Glorot, X. and Bengio, Y. (2010). Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, pp. 249\u2013256."},{"key":"S135132492000008X_ref3","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-60042-0_66"},{"key":"S135132492000008X_ref47","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1146"},{"key":"S135132492000008X_ref38","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"S135132492000008X_ref25","unstructured":"Gulcehre, C. , Moczulski, M. , Denil, M. and Bengio, Y. (2016) Noisy activation functions. In International Conference on Machine Learning, pp. 3059\u20133068."},{"key":"S135132492000008X_ref2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-70096-0_51"},{"key":"S135132492000008X_ref36","unstructured":"Mourad, A. and Darwish, K. (2013). Subjectivity and sentiment analysis of modern standard Arabic and Arabic microblogs. In Proceedings of the 4th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pp. 55\u201364."},{"key":"S135132492000008X_ref16","unstructured":"Dahou, A. , Xiong, S. , Zhou, J. , Haddoud, M.H. and Duan, P. (2016). Word embeddings and convolutional neural network for Arabic sentiment classification. In COLING 2016, pp. 2418\u20132427."},{"key":"S135132492000008X_ref52","doi-asserted-by":"publisher","DOI":"10.1162\/COLI_a_00169"},{"key":"S135132492000008X_ref1","doi-asserted-by":"publisher","DOI":"10.1109\/AEECT.2013.6716448"},{"key":"S135132492000008X_ref39","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2016.07.001"},{"key":"S135132492000008X_ref14","first-page":"2493","article-title":"Natural language processing (almost) from scratch","volume":"12","author":"Collobert","year":"2011","journal-title":"Journal of Machine Learning Research"},{"key":"S135132492000008X_ref46","unstructured":"Socher, R. , Perelygin, A. , Wu, J. , Chuang, J. , Manning, C.D. , Ng, A. and Potts, C. (2013). Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1631\u20131642."},{"key":"S135132492000008X_ref49","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"van der Maaten","year":"2008","journal-title":"Journal of Machine Learning Research"},{"key":"S135132492000008X_ref19","doi-asserted-by":"publisher","DOI":"10.1109\/CloudTech.2017.8284706"},{"key":"S135132492000008X_ref42","doi-asserted-by":"publisher","DOI":"10.1002\/asi.21598"},{"key":"S135132492000008X_ref5","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W15-3202"},{"key":"S135132492000008X_ref37","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1299"},{"key":"S135132492000008X_ref44","doi-asserted-by":"crossref","unstructured":"Shen, D. , Wang, G. , Wang, W. , Min, M.R. , Su, Q. , Zhang, Y. , Li, C. , Henao, R. and Carin, L. (2018). Baseline needs more love: On simple word-embedding-based models and associated pooling mechanisms. arXiv preprint arXiv:1805.09843.","DOI":"10.18653\/v1\/P18-1041"},{"key":"S135132492000008X_ref41","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S17-2088"},{"key":"S135132492000008X_ref24","doi-asserted-by":"crossref","unstructured":"Gridach, M. , Haddad, H. and Mulki, H. (2017). Empirical evaluation of word representations on Arabic sentiment analysis. In International Conference on Arabic Language Processing (ICALP), pp. 147\u2013158.","DOI":"10.1007\/978-3-319-73500-9_11"},{"key":"S135132492000008X_ref8","unstructured":"Aly, M. and Atiya, A. (2013). LABR: A large scale Arabic book reviews dataset. In ACL (2), pp. 494\u2013498."},{"key":"S135132492000008X_ref53","unstructured":"Zbib, R. , Malchiodi, E. , Devlin, J. , Stallard, D. , Matsoukas, S. , Schwartz, R. , Makhoul, J. , Zaidan, O.F. and Callison-Burch, C. (2012). Machine translation of Arabic dialects. In Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 49\u201359."},{"key":"S135132492000008X_ref33","unstructured":"Mikolov, T. , Sutskever, I. , Chen, K. , Corrado, G.S. and Dean, J. (2013) Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems, pp. 3111\u20133119."},{"key":"S135132492000008X_ref18","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S17-2133"},{"key":"S135132492000008X_ref17","first-page":"2121","article-title":"Adaptive subgradient methods for online learning and stochastic optimization","volume":"12","author":"Duchi","year":"2011","journal-title":"Journal of Machine Learning Research"},{"key":"S135132492000008X_ref50","doi-asserted-by":"publisher","DOI":"10.1145\/2838931.2838932"},{"key":"S135132492000008X_ref21","unstructured":"Firth, J.R. (1957). A synopsis of linguistic theory 1930\u20131955. In Studies in linguistic analysis (pp. 1\u201332). Oxford: Philological Society. [Reprinted in F. R. Palmer (Ed.) (1968). Selected papers of J. R. Firth 1952\u20131959. London: Longman.]"},{"key":"S135132492000008X_ref13","unstructured":"Chiang, D. , Diab, M. , Habash, N. , Rambow, O. and Shareef, S. (2006). Parsing arabic dialects. In 11th Conference of the European Chapter of the Association for Computational Linguistics."},{"key":"S135132492000008X_ref48","unstructured":"Tieleman, T. and Hinton, G. (2012). Lecture 6.5\u2013RmsProp: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural Networks for Machine Learning."},{"key":"S135132492000008X_ref20","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-18117-2_2"},{"key":"S135132492000008X_ref35","unstructured":"Mohammad, S.M. , Kiritchenko, S. and Zhu, X. (2013). NRC-Canada: Building the state-of-the-art in sentiment analysis of tweets. arXiv preprint arXiv:1308.6242."},{"key":"S135132492000008X_ref54","unstructured":"Zeiler, M.D. (2012). ADADELTA: An adaptive learning rate method. arXiv preprint arXiv:1212.5701."},{"key":"S135132492000008X_ref11","unstructured":"Baniata, L.H. and Park, S.-B. (2016). Sentence representation network for Arabic sentiment analysis. In Proceedings of the Korean Information Science Society, pp. 470\u2013472."},{"key":"S135132492000008X_ref51","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-18111-0_32"},{"key":"S135132492000008X_ref32","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-1307"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S135132492000008X","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,5,31]],"date-time":"2021-05-31T10:04:30Z","timestamp":1622455470000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S135132492000008X\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,16]]},"references-count":55,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,5]]}},"alternative-id":["S135132492000008X"],"URL":"https:\/\/doi.org\/10.1017\/s135132492000008x","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,3,16]]},"assertion":[{"value":"\u00a9 Cambridge University Press 2020","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}}]}}