{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,16]],"date-time":"2026-01-16T06:52:04Z","timestamp":1768546324147,"version":"3.49.0"},"reference-count":39,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2023,3,22]],"date-time":"2023-03-22T00:00:00Z","timestamp":1679443200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:p>Deep neural networks have been proven effective in classifying human interactions into emotions, especially by encoding multiple input modalities. In this work, we assess the robustness of a transformer-based multimodal audio-text classifier for emotion recognition, by perturbing the input at inference time using attacks which we design specifically to corrupt information deemed important for emotion recognition. To measure the impact of the attacks on the classifier, we compare between the accuracy of the classifier on the perturbed input and on the original, unperturbed input. Our results show that the multimodal classifier is more resilient to perturbation attacks than the equivalent unimodal classifiers, suggesting that the two modalities are encoded in a way that allows the classifier to benefit from one modality even when the other one is slightly damaged.<\/jats:p>","DOI":"10.3389\/frai.2023.1091443","type":"journal-article","created":{"date-parts":[[2023,3,22]],"date-time":"2023-03-22T13:32:38Z","timestamp":1679491958000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Masking important information to assess the robustness of a multimodal classifier for emotion recognition"],"prefix":"10.3389","volume":"6","author":[{"given":"Dror","family":"Cohen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ido","family":"Rosenberger","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Moshe","family":"Butman","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kfir","family":"Bar","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2023,3,22]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"e17","DOI":"10.1017\/ATSIP.2020.14","article-title":"Dimensional speech emotion recognition from speech features and word embeddings by using multitask learning","volume":"9","author":"Atmaja","year":"2020","journal-title":"APSIPA Trans. Signal Inform. Process."},{"key":"B2","doi-asserted-by":"crossref","first-page":"1081","DOI":"10.1109\/TENCON50793.2020.9293899","article-title":"\u201cPredicting valence and arousal by aggregating acoustic features for acoustic-linguistic information fusion,\u201d","author":"Atmaja","year":"2020","journal-title":"2020 IEEE Region 10 Conference (TENCON)"},{"key":"B3","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1016\/j.specom.2022.03.002","article-title":"Survey on bimodal speech emotion recognition from acoustic and linguistic information fusion","volume":"140","author":"Atmaja","year":"2022","journal-title":"Speech Commun"},{"key":"B4","first-page":"12449","article-title":"\u201cwav2vec 2.0: a framework for self-supervised learning of speech representations,\u201d","author":"Baevski","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B5","author":"Bolinger","year":"1986","journal-title":"Intonation and Its Parts: Melody in Spoken English"},{"key":"B6","doi-asserted-by":"publisher","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: interactive emotional dyadic motion capture database","volume":"42","author":"Busso","year":"2008","journal-title":"Lang. Resour. Eval."},{"key":"B7","doi-asserted-by":"crossref","first-page":"205","DOI":"10.1145\/1027933.1027968","article-title":"\u201cAnalysis of emotion recognition using facial expressions, speech and multimodal information,\u201d","author":"Busso","year":"2004","journal-title":"Proceedings of the 6th International Conference on Multimodal Interfaces"},{"key":"B8","doi-asserted-by":"publisher","first-page":"2593036","DOI":"10.1155\/2019\/2593036","article-title":"Audio-textual emotion recognition based on improved neural networks","volume":"2019","author":"Cai","year":"2019","journal-title":"Math. Prob. Eng."},{"key":"B9","first-page":"374","article-title":"\u201cA multi-scale fusion framework for bimodal speech emotion recognition,\u201d","author":"Chen","year":"2020","journal-title":"Interspeech"},{"key":"B10","doi-asserted-by":"publisher","first-page":"247","DOI":"10.21437\/Interspeech.2018-2466","article-title":"\u201cDeep neural networks for emotion recognition combining audio and transcripts,\u201d","author":"Cho","year":"2019","journal-title":"Proceedings of Interspeech 2018"},{"key":"B11","doi-asserted-by":"publisher","first-page":"4171","DOI":"10.18653\/v1\/N19-1423","article-title":"\u201cBERT: Pre-training of deep bidirectional transformers for language understanding,\u201d","author":"Devlin","year":"2019","journal-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)"},{"key":"B12","first-page":"745","article-title":"\u201cFacial expression recognition via deep learning,\u201d","author":"Fathallah","year":"2017","journal-title":"2017 IEEE\/ACS 14th International Conference on Computer Systems and Applications (AICCSA)"},{"key":"B13","first-page":"4243","article-title":"\u201cMultimodal emotion recognition using cross-modal attention and 1D convolutional neural networks,\u201d","author":"Krishna","year":"2020","journal-title":"Interspeech"},{"key":"B14","article-title":"Foundations and recent trends in multimodal machine learning: principles, challenges, and open questions","author":"Liang","year":"2022","journal-title":"arXiv preprint arXiv:2209.03430"},{"key":"B15","first-page":"379","article-title":"\u201cGroup gated fusion on attention-based bidirectional alignment for multimodal emotion recognition,\u201d","volume-title":"Proceedings of Interspeech 2020","author":"Liu","year":"2022"},{"key":"B16","doi-asserted-by":"crossref","first-page":"18","DOI":"10.25080\/Majora-7b98e3ed-003","article-title":"\u201clibrosa: audio and music signal analysis in Python,\u201d","author":"McFee","year":"2015","journal-title":"Proceedings of the 14th Python in Science Conference"},{"key":"B17","doi-asserted-by":"crossref","first-page":"2227","DOI":"10.1109\/ICASSP.2017.7952552","article-title":"\u201cAutomatic speech emotion recognition using recurrent neural networks with local attention,\u201d","author":"Mirsamadi","year":"2017","journal-title":"2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)"},{"key":"B18","doi-asserted-by":"crossref","first-page":"7390","DOI":"10.1109\/ICASSP.2019.8682541","article-title":"\u201cImproving speech emotion recognition with unsupervised representation learning on unlabeled speech,\u201d","author":"Neumann","year":"2019","journal-title":"ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)"},{"key":"B19","doi-asserted-by":"publisher","first-page":"3400","DOI":"10.21437\/Interspeech.2021-703","article-title":"\u201cEmotion recognition from speech using wav2vec 2.0 embeddings,\u201d","author":"Pepino","year":"2021","journal-title":"Proceedings of Interspeech 2021"},{"key":"B20","doi-asserted-by":"publisher","first-page":"320","DOI":"10.1075\/gest.14.3.03per","article-title":"Iconicity in vocalization, comparisons with gesture, and implications for theories on the evolution of language","volume":"14","author":"Perlman","year":"2014","journal-title":"Gesture"},{"key":"B21","doi-asserted-by":"crossref","first-page":"236","DOI":"10.1142\/9789814603638_0030","article-title":"\u201cIterative vocal charades: the emergence of conventions in vocal communication,\u201d","author":"Perlman","year":"2014","journal-title":"Evolution of Language: Proceedings of the 10th International Conference (EVOLANG10)"},{"key":"B22","doi-asserted-by":"publisher","first-page":"20181634","DOI":"10.1098\/rspb.2018.1634","article-title":"Voice pitch modulation in human mate choice","volume":"285","author":"Pisanski","year":"2018","journal-title":"Proc. R. Soc. B"},{"key":"B23","doi-asserted-by":"crossref","first-page":"439","DOI":"10.1109\/ICDM.2016.0055","article-title":"\u201cConvolutional MKL based multimodal emotion recognition and sentiment analysis,\u201d","author":"Poria","year":"2016","journal-title":"2016 IEEE 16th International Conference on Data Mining (ICDM)"},{"key":"B24","doi-asserted-by":"publisher","first-page":"17","DOI":"10.1109\/MIS.2018.2882362","article-title":"Multimodal sentiment analysis: addressing key issues and setting up the baselines","volume":"33","author":"Poria","year":"2018","journal-title":"IEEE Intell. Syst."},{"key":"B25","doi-asserted-by":"crossref","first-page":"75","DOI":"10.1145\/2988257.2988268","article-title":"\u201cMultimodal emotion recognition for avec 2016 challenge,\u201d","author":"Povolny","year":"2016","journal-title":"Proceedings of the 6th International Workshop on Audio\/Visual Emotion Challenge"},{"key":"B26","first-page":"1","article-title":"\u201cMultimodal emotion recognition using deep learning architectures,\u201d","author":"Ranganathan","year":"2016","journal-title":"2016 IEEE Winter Conference on Applications of Computer Vision (WACV)"},{"key":"B27","first-page":"1","article-title":"\u201cIntroducing the recola multimodal corpus of remote collaborative and affective interactions,\u201d","author":"Ringeval","year":"2013","journal-title":"2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG)"},{"key":"B28","doi-asserted-by":"publisher","first-page":"646","DOI":"10.3390\/e21070646","article-title":"Emotion recognition from skeletal movements","volume":"21","author":"Sapi\u0144ski","year":"2019","journal-title":"Entropy"},{"key":"B29","article-title":"\u201cRobustness analysis of video-language models against visual and language perturbations,\u201d","author":"Schiappa","year":"2022","journal-title":"Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track"},{"key":"B30","article-title":"Bias and fairness on multimodal emotion detection algorithms","author":"Schmitz","year":"2022","journal-title":"arXiv preprint arXiv:2205.08383"},{"key":"B31","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.patrec.2021.03.007","article-title":"Leveraging recent advances in deep learning for audio-visual emotion recognition","volume":"146","author":"Schoneveld","year":"2021","journal-title":"Pattern Recogn. Lett."},{"key":"B32","doi-asserted-by":"crossref","first-page":"5149","DOI":"10.1109\/ICASSP.2012.6289079","article-title":"\u201cJapanese and Korean voice search,\u201d","author":"Schuster","year":"2012","journal-title":"2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)"},{"key":"B33","doi-asserted-by":"publisher","first-page":"167","DOI":"10.1016\/j.jml.2006.03.002","article-title":"Analog acoustic expression in speech communication","volume":"55","author":"Shintel","year":"2006","journal-title":"J. Mem. Lang."},{"key":"B34","article-title":"Analyzing the influence of dataset composition for emotion recognition","author":"Sutherland","year":"2021","journal-title":"arXiv preprint arXiv:2103.03700"},{"key":"B35","article-title":"Multi-modal emotion recognition on IEMOCAP dataset using deep learning","author":"Tripathi","year":"2018","journal-title":"arXiv preprint arXiv:1804.05788"},{"key":"B36","unstructured":"\u201cAttention is all you need,\u201d\n            VaswaniA.\n            ShazeerN.\n            ParmarN.\n            UszkoreitJ.\n            JonesL.\n            GomezA. N.\n          Advances in Neural Information Processing Systems2017"},{"key":"B37","doi-asserted-by":"publisher","first-page":"38","DOI":"10.18653\/v1\/2020.emnlp-demos.6","article-title":"\u201cTransformers: State-of-the-art natural language processing,\u201d","author":"Wolf","year":"2019","journal-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations"},{"key":"B38","first-page":"3340","article-title":"\u201cDefending multimodal fusion models against single-source adversaries,\u201d","author":"Yang","year":"2021","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B39","doi-asserted-by":"publisher","first-page":"3030","DOI":"10.1109\/TCSVT.2017.2719043","article-title":"Learning affective features with a hybrid deep model for audio\u2013visual emotion recognition","volume":"28","author":"Zhang","year":"2017","journal-title":"IEEE Trans. Circ. Syst. Video Technol."}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2023.1091443\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,22]],"date-time":"2023-03-22T13:33:04Z","timestamp":1679491984000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2023.1091443\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,22]]},"references-count":39,"alternative-id":["10.3389\/frai.2023.1091443"],"URL":"https:\/\/doi.org\/10.3389\/frai.2023.1091443","relation":{},"ISSN":["2624-8212"],"issn-type":[{"value":"2624-8212","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,22]]},"article-number":"1091443"}}