{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,16]],"date-time":"2026-01-16T18:14:34Z","timestamp":1768587274014,"version":"3.49.0"},"reference-count":39,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2021,4,12]],"date-time":"2021-04-12T00:00:00Z","timestamp":1618185600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The paper proposes three modeling techniques to improve the performance evaluation of the call center agent. The first technique is speech processing supported by an attention layer for the agent\u2019s recorded calls. The speech comprises 65 features for the ultimate determination of the context of the call using the Open-Smile toolkit. The second technique uses the Max Weights Similarity (MWS) approach instead of the Softmax function in the attention layer to improve the classification accuracy. MWS function replaces the Softmax function for fine-tuning the output of the attention layer for processing text. It is formed by determining the similarity in the distance of input weights of the attention layer to the weights of the max vectors. The third technique combines the agent\u2019s recorded call speech with the corresponding transcribed text for binary classification. The speech modeling and text modeling are based on combinations of the Convolutional Neural Networks (CNNs) and Bi-directional Long-Short Term Memory (BiLSTMs). In this paper, the classification results for each model (text versus speech) are proposed and compared with the multimodal approach\u2019s results. The multimodal classification provided an improvement of (0.22%) compared with acoustic model and (1.7%) compared with text model.<\/jats:p>","DOI":"10.3390\/s21082720","type":"journal-article","created":{"date-parts":[[2021,4,12]],"date-time":"2021-04-12T21:47:33Z","timestamp":1618264053000},"page":"2720","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["A Multimodal Approach to Improve Performance Evaluation of Call Center Agent"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1024-5469","authenticated-orcid":false,"given":"Abdelrahman","family":"Ahmed","sequence":"first","affiliation":[{"name":"Department of Electronic Engineering, University of Seville, 41092 Seville, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0823-8390","authenticated-orcid":false,"given":"Khaled","family":"Shaalan","sequence":"additional","affiliation":[{"name":"Faculty of Engineering &amp; IT, British University in Dubai, Dubai 345015, United Arab Emirates"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2612-0388","authenticated-orcid":false,"given":"Sergio","family":"Toral","sequence":"additional","affiliation":[{"name":"Department of Electronic Engineering, University of Seville, 41092 Seville, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5516-7225","authenticated-orcid":false,"given":"Yasser","family":"Hifny","sequence":"additional","affiliation":[{"name":"Faculty of Computers and Artificial Intelligence, University of Helwan, Helwan 11795, Egypt"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,4,12]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"141","DOI":"10.1007\/s11846-011-0076-3","article-title":"Social ties and subjective performance evaluations: An empirical investigation","volume":"7","author":"Breuer","year":"2013","journal-title":"Rev. Manag. Sci."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.4102\/sajhrm.v16i0.905","article-title":"Exploring employee retention and intention to leave within a call center","volume":"16","author":"Dhanpat","year":"2018","journal-title":"SA J. Hum. Resour. Manag."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"408","DOI":"10.1016\/j.jebo.2016.12.016","article-title":"Subjective performance evaluations and employee careers","volume":"134","author":"Frederiksen","year":"2017","journal-title":"J. Econ. Behav. Organ."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"797","DOI":"10.1016\/j.indmarman.2005.01.002","article-title":"Cultural vs. operational market orientation and objective vs. subjective performance: Perspective of production and operations","volume":"34","year":"2005","journal-title":"Ind. Mark. Manag."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1080\/08911762.2018.1427293","article-title":"Emotional Exhaustion in Offshore Call Centers: A Comparative Study","volume":"32","author":"Echchakoui","year":"2019","journal-title":"J. Glob. Mark."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ahmed, A., Hifny, Y., Toral, S., and Shaalan, K. (2018). A Call Center Agent Productivity Modeling Using Discriminative Approaches. Intelligent Natural Language Processing: Trends and Applications, Springer. Book Section 1.","DOI":"10.1007\/978-3-319-67056-0_24"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Ahmed, A., Toral, S., and Shaalan, K. (2016). Agent productivity measurement in call center using machine learning. International Conference on Advanced Intelligent Systems and Informatics, Springer.","DOI":"10.1007\/978-3-319-48308-5_16"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1142\/9789813229396_0011","article-title":"End-to-End Lexicon Free Arabic Speech Recognition Using Recurrent Neural Networks","volume":"4","author":"Ahmed","year":"2018","journal-title":"Comput. Linguist. Speech Image Process. Arab. Lang."},{"key":"ref_9","first-page":"1","article-title":"Feature extraction methods LPC, PLP and MFCC in speech recognition","volume":"1","author":"Dave","year":"2013","journal-title":"Int. J. Adv. Res. Eng. Technol."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1016\/j.eswa.2004.08.008","article-title":"A web-based system for analyzing the voices of call center customers in the service industry","volume":"28","author":"Bae","year":"2005","journal-title":"Expert Syst. Appl."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Karakus, B., and Aydin, G. (2016, January 11\u201313). Call center performance evaluation using big data analytics. Proceedings of the 2016 International Symposium on Networks, Computers and Communications (ISNCC), Hammamet, Tunisia.","DOI":"10.1109\/ISNCC.2016.7746116"},{"key":"ref_12","first-page":"1","article-title":"Automatic Evaluation Software for Contact Centre Agents\u2019 voice Handling Performance","volume":"5","author":"Perera","year":"2019","journal-title":"Int. J. Sci. Res. Publ."},{"key":"ref_13","first-page":"133","article-title":"Voice call analytics using natural language processing","volume":"4","author":"Sudarsan","year":"2019","journal-title":"Int. J. Stat. Appl. Math."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Ahmed, A., Hifny, Y., Shaalan, K., and Toral, S. (2016). Lexicon free Arabic speech recognition recipe. International Conference on Advanced Intelligent Systems and Informatics, Springer.","DOI":"10.1007\/978-3-319-48308-5_15"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Neumann, M., and Vu, N.T. (2017). Attentive convolutional neural network based speech emotion recognition: A study on the impact of input features, signal length, and acted speech. arXiv.","DOI":"10.21437\/Interspeech.2017-917"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Hifny, Y., and Ali, A. (2019, January 12\u201317). Efficient Arabic Emotion Recognition Using Deep Neural Networks. Proceedings of the ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683632"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Cho, J., Pappagari, R., Kulkarni, P., Villalba, J., Carmiel, Y., and Dehak, N. (2019). Deep neural networks for emotion recognition combining audio and transcripts. arXiv.","DOI":"10.21437\/Interspeech.2018-2466"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Li, P., Jiang, Z., Yin, S., Song, D., Ouyang, P., Liu, L., and Wei, S. (2020, January 4\u20139). PAGAN: A Phase-Adapted Generative Adversarial Networks for Speech Enhancement. Proceedings of the ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054256"},{"key":"ref_19","unstructured":"Cleveland, B. (2012). Call Center Management on Fast Forward: Succeeding in the New Era of Customer Relationships, ICMI Press."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1080\/19312450709336664","article-title":"Answering the call for a standard reliability measure for coding data","volume":"1","author":"Hayes","year":"2007","journal-title":"Commun. Methods Meas."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Trigeorgis, G., Ringeval, F., Brueckner, R., Marchi, E., Nicolaou, M.A., Schuller, B., and Zafeiriou, S. (2016, January 20\u201325). Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network. Proceedings of the 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472669"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"182","DOI":"10.1016\/j.neucom.2020.07.027","article-title":"Character-level neural network model based on Nadam optimization and its application in clinical concept extraction","volume":"414","author":"Li","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_23","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 8\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Los Angeles, CA, USA."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1016\/j.sigpro.2017.12.008","article-title":"Recurrent attention network using spatial-temporal relations for action recognition","volume":"145","author":"Zhang","year":"2018","journal-title":"Signal Process."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Ahmed, A., Toral, S., Shaalan, K., and Hifny, Y. (2020). Agent Productivity Modeling in a Call Center Domain Using Attentive Convolutional Neural Networks. Sensors, 20.","DOI":"10.3390\/s20195489"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Eyben, F., W\u00f6llmer, M., and Schuller, B. (2010, January 21\u201325). Opensmile: The munich versatile and fast open-source audio feature extractor. Proceedings of the 18th ACM international conference on Multimedia, Nice, France.","DOI":"10.1145\/1873951.1874246"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Palaz, D., and Collobert, R. (2015). Analysis of CNN-Based Speech Recognition System Using Raw Speech as Input, Idiap. Technical Report.","DOI":"10.21437\/Interspeech.2015-3"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Norouzian, A., Mazoure, B., Connolly, D., and Willett, D. (2019, January 12\u201317). Exploring attention mechanism for acoustic-based classification of speech utterances into system-directed and non-system-directed. Proceedings of the ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683565"},{"key":"ref_29","unstructured":"Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y.N. (2017, January 17). Convolutional sequence to sequence learning. Proceedings of the 34th International Conference on Machine Learning-Volume 70, Sydney, Australia."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Bridle, J.S. (1990). Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition. Neurocomputing, Springer.","DOI":"10.1007\/978-3-642-76153-9_28"},{"key":"ref_31","unstructured":"Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A.Y. (July, January 28). Multimodal deep learning. Proceedings of the ICML, Bellevue, WA, USA."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1109\/MSP.2017.2738401","article-title":"Deep multimodal learning: A survey on recent advances and trends","volume":"34","author":"Ramachandram","year":"2017","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1007\/s12193-015-0195-2","article-title":"Emonets: Multimodal deep learning approaches for emotion recognition in video","volume":"10","author":"Kahou","year":"2016","journal-title":"J. Multimodal User Interfaces"},{"key":"ref_34","first-page":"423","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Ahuja","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: Interactive emotional dyadic motion capture database","volume":"42","author":"Busso","year":"2008","journal-title":"Lang. Resour. Eval."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Broux, P.-A., Desnous, F., Larcher, A., Petitrenaud, S., Carrive, J., and Meignier, S. (2018). S4D: Speaker Diarization Toolkit in Python. Interspeech.","DOI":"10.21437\/Interspeech.2018-1232"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Schuller, B., Steidl, S., Batliner, A., Hirschberg, J., Burgoon, J.K., Baird, A., Elkins, A., Zhang, Y., Coutinho, E., and Evanini, K. (2016, January 8\u201312). The interspeech 2016 computational paralinguistics challenge: Deception, sincerity & native language. Proceedings of the 17th Annual Conference of the International Speech Communication Association (Interspeech 2016), San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-129"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_39","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/8\/2720\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T13:35:52Z","timestamp":1760362552000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/8\/2720"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,12]]},"references-count":39,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2021,4]]}},"alternative-id":["s21082720"],"URL":"https:\/\/doi.org\/10.3390\/s21082720","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,4,12]]}}}