{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T20:37:24Z","timestamp":1782506244864,"version":"3.54.5"},"reference-count":64,"publisher":"Springer Science and Business Media LLC","issue":"9","license":[{"start":{"date-parts":[[2024,8,2]],"date-time":"2024-08-02T00:00:00Z","timestamp":1722556800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,8,2]],"date-time":"2024-08-02T00:00:00Z","timestamp":1722556800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Speech classification tasks often require powerful language understanding models to grasp useful features, which becomes problematic when limited training data is available. To attain superior classification performance, we propose to harness the inherent value of multimodal representations by transcribing speech using automatic speech recognition models and translating the transcripts into different languages via pretrained translation models. We thus obtain an audio\u2013textual (multimodal) representation for each data sample. Subsequently, we combine language-specific Bidirectional Encoder Representations from Transformers with Wav2Vec2.0 audio features via a novel cascaded cross-modal transformer (CCMT). Our model is based on two cascaded transformer blocks. The first one combines text-specific features from distinct languages, while the second one combines acoustic features with multilingual features previously learned by the first transformer block. We employed our system in the Requests Sub-Challenge of the ACM Multimedia 2023 Computational Paralinguistics Challenge. CCMT was declared the winning solution, obtaining an unweighted average recall of 65.41% and 85.87% for complaint and request detection, respectively. Moreover, we applied our framework on the Speech Commands v2 and HVB dialog data sets, surpassing previous studies reporting results on these benchmarks. Our code is freely available for download at: <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/ristea\/ccmt\">https:\/\/github.com\/ristea\/ccmt<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s10462-024-10869-1","type":"journal-article","created":{"date-parts":[[2024,8,2]],"date-time":"2024-08-02T13:34:19Z","timestamp":1722605659000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Cascaded cross-modal transformer for audio\u2013textual classification"],"prefix":"10.1007","volume":"57","author":[{"given":"Nicolae-C\u0103t\u0103lin","family":"Ristea","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrei","family":"Anghel","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Radu Tudor","family":"Ionescu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,8,2]]},"reference":[{"key":"10869_CR1","doi-asserted-by":"publisher","first-page":"204","DOI":"10.1016\/j.inffus.2021.06.003","volume":"76","author":"SA Abdu","year":"2021","unstructured":"Abdu SA, Yousef AH, Salem A (2021) Multimodal video sentiment analysis using deep learning approaches, a survey. Inf Fusion 76:204\u2013226","journal-title":"Inf Fusion"},{"key":"10869_CR2","unstructured":"Akbari H, Yuan L, Qian R, Chuang W-H, Chang S-F, Cui Y, Gong B (2021) VATT: transformers for multimodal self-supervised learning from raw video, audio and text. In: Proceedings of NeurIPS, vol 34, pp 24206\u201324221"},{"key":"10869_CR3","unstructured":"Baevski A, Zhou Y, Mohamed A, Auli M (2020) wav2vec 2.0: a framework for self-supervised learning of speech representations. In: Proceedings of NeurIPS, vol 33, pp 12449\u201312460"},{"key":"10869_CR4","doi-asserted-by":"publisher","first-page":"635","DOI":"10.1016\/j.procs.2015.02.112","volume":"46","author":"J Bhaskar","year":"2015","unstructured":"Bhaskar J, Sruthi K, Nedungadi P (2015) Hybrid approach for emotion classification of audio conversation based on text and speech mining. Procedia Comput Sci 46:635\u2013643","journal-title":"Procedia Comput Sci"},{"issue":"6","key":"10869_CR5","doi-asserted-by":"publisher","first-page":"121","DOI":"10.1007\/s00138-021-01249-8","volume":"32","author":"SY Boulahia","year":"2021","unstructured":"Boulahia SY, Amamra A, Madi MR, Daikh S (2021) Early, intermediate and late fusion strategies for robust deep learning-based multimodal action recognition. Mach Vis Appl 32(6):121","journal-title":"Mach Vis Appl"},{"key":"10869_CR6","doi-asserted-by":"publisher","first-page":"722","DOI":"10.1109\/LSP.2022.3151551","volume":"29","author":"N Braunschweiler","year":"2022","unstructured":"Braunschweiler N, Doddipatla R, Keizer S, Stoyanchev S (2022) Factors in emotion recognition with deep learning models using speech and text on multiple corpora. IEEE Signal Process Lett 29:722\u2013726","journal-title":"IEEE Signal Process Lett"},{"key":"10869_CR7","unstructured":"Ca\u00f1ete J, Chaperon G, Fuentes R, Ho J-H, Kang H, P\u00e9rez J (2020) Spanish pre-trained BERT model and evaluation data. In: Proceedings of PML4DC (ICLR workshop)"},{"key":"10869_CR8","unstructured":"Chung HW, Hou L, Longpre S, Zoph B, Tay Y, Fedus W, Li E, Wang X, Dehghani M, Brahma S et al (2022) Scaling instruction-finetuned language models. arXiv preprint. arXiv:2210.11416"},{"key":"10869_CR9","doi-asserted-by":"publisher","DOI":"10.1145\/3586075","author":"R Das","year":"2023","unstructured":"Das R, Singh TD (2023) Multimodal sentiment analysis: a survey of methods, trends and challenges. ACM Comput Surv. https:\/\/doi.org\/10.1145\/3586075","journal-title":"ACM Comput Surv"},{"key":"10869_CR10","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova LK (2019) BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT, pp 4171\u20134186"},{"key":"10869_CR11","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, Uszkoreit J, Houlsby N (2021) An image is worth $$16\\times 16$$ words: transformers for image recognition at scale. In: Proceedings of ICLR"},{"key":"10869_CR12","doi-asserted-by":"crossref","unstructured":"Dumitrescu \u015eD, Avram A-M, Pyysalo S (2020) The birth of Romanian BERT. In: Proceedings of EMNLP, pp 4324\u20134328","DOI":"10.18653\/v1\/2020.findings-emnlp.387"},{"issue":"5","key":"10869_CR13","doi-asserted-by":"publisher","first-page":"829","DOI":"10.1162\/neco_a_01273","volume":"32","author":"J Gao","year":"2020","unstructured":"Gao J, Li P, Chen Z, Zhang J (2020) A survey on deep learning for multimodal data fusion. Neural Comput 32(5):829\u2013864","journal-title":"Neural Comput"},{"issue":"2","key":"10869_CR14","doi-asserted-by":"publisher","first-page":"83","DOI":"10.3390\/info13020083","volume":"13","author":"A Gasparetto","year":"2022","unstructured":"Gasparetto A, Marcuzzo M, Zangari A, Albarelli A (2022) A survey on text classification algorithms: from text to predictions. Information 13(2):83","journal-title":"Information"},{"key":"10869_CR15","doi-asserted-by":"crossref","unstructured":"Gemmeke JF, Ellis DPW, Freedman D, Jansen A, Lawrence W, Moore RC, Plakal M, Ritter M (2017) Audio Set: an ontology and human-labeled dataset for audio events. In: Proceedings of ICASSP. IEEE, pp 776\u2013780","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"10869_CR16","doi-asserted-by":"crossref","unstructured":"Georgescu M-I, Fonseca E, Ionescu RT, Lucic M, Schmid C, Arnab A (2023) Audiovisual masked autoencoders. In: Proceedings of ICCV 16144\u201316154","DOI":"10.1109\/ICCV51070.2023.01479"},{"key":"10869_CR17","doi-asserted-by":"crossref","unstructured":"Gong Y, Chung Y-A, Glass J (2021) AST: audio spectrogram transformer. In: Proceedings of INTERSPEECH, pp 571\u2013575","DOI":"10.21437\/Interspeech.2021-698"},{"key":"10869_CR18","doi-asserted-by":"crossref","unstructured":"Gong Y, Lai C-I, Chung Y-A, Glass J (2022) SSAST: self-supervised audio spectrogram transformer. In: Proceedings of AAAI, vol 36, pp 10699\u201310709","DOI":"10.1609\/aaai.v36i10.21315"},{"key":"10869_CR19","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of CVPR, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"issue":"8","key":"10869_CR20","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735\u20131780","journal-title":"Neural Comput"},{"issue":"1","key":"10869_CR21","doi-asserted-by":"publisher","first-page":"136","DOI":"10.1038\/s41746-020-00341-z","volume":"3","author":"S-C Huang","year":"2020","unstructured":"Huang S-C, Pareek A, Seyyedi S, Banerjee I, Lungren MP (2020a) Fusion of medical imaging and electronic health records using deep learning: a systematic review and implementation guidelines. NPJ Digit Med 3(1):136","journal-title":"NPJ Digit Med"},{"key":"10869_CR22","doi-asserted-by":"crossref","unstructured":"Huang J, Tao J, Liu B, Lian Z, Niu M (2020b) Multimodal transformer fusion for continuous emotion recognition. In Proceedings of ICASSP. IEEE, pp 3507\u20133511","DOI":"10.1109\/ICASSP40776.2020.9053762"},{"key":"10869_CR23","unstructured":"Huang P-Y, Xu H, Li J, Baevski A, Auli M, Galuba W, Metze F, Feichtenhofer C (2022) Masked autoencoders that listen. In: Proceedings of NeurIPS, vol 35, pp 28708\u201328720"},{"issue":"2s","key":"10869_CR24","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3545572","volume":"19","author":"S Jabeen","year":"2023","unstructured":"Jabeen S, Li X, Amin MS, Bourahla O, Li S, Jabbar A (2023) A review on methods and applications in multimodal deep learning. ACM Trans Multimed Comput Commun Appl 19(2s):1\u201341","journal-title":"ACM Trans Multimed Comput Commun Appl"},{"issue":"6","key":"10869_CR25","doi-asserted-by":"publisher","first-page":"2891","DOI":"10.3390\/app12062891","volume":"12","author":"M Khadhraoui","year":"2022","unstructured":"Khadhraoui M, Bellaaj H, Ammar MB, Hamam H, Jmaiel M (2022) Survey of BERT-base models for scientific text classification: COVID-19 case study. Appl Sci 12(6):2891","journal-title":"Appl Sci"},{"key":"10869_CR26","unstructured":"Kingma DP, Ba J (2014) ADAM: a method for stochastic optimization. In: Proceedings of ICLR"},{"key":"10869_CR27","doi-asserted-by":"publisher","first-page":"2880","DOI":"10.1109\/TASLP.2020.3030497","volume":"28","author":"Q Kong","year":"2020","unstructured":"Kong Q, Cao Y, Iqbal T, Wang Y, Wang W, Plumbley MD (2020) PANNs: large-scale pretrained audio neural networks for audio pattern recognition. IEEE\/ACM Trans Audio Speech Lang Process 28:2880\u20132894","journal-title":"IEEE\/ACM Trans Audio Speech Lang Process"},{"key":"10869_CR28","unstructured":"Lackovic N, Montaci\u00e9 C, Lalande G, Caraty M-J (2022) Prediction of user request and complaint in spoken customer-agent conversations. arXiv preprint. arXiv:2208.10249"},{"key":"10869_CR29","unstructured":"Le H, Vial L, Frej J, Segonne V, Coavoux M, Lecouteux B, Allauzen A, Crabb\u00e9 B, Besacier L, Schwab D (2020) FlauBERT: unsupervised language model pre-training for French. In: Proceedings of LREC, pp 2479\u20132490"},{"key":"10869_CR30","doi-asserted-by":"crossref","unstructured":"Lee W-Y, Jovanov L, Philips W (2022) Cross-modality attention and multimodal fusion transformer for pedestrian detection. In: Proceedings of ECCV. Springer, Cham, pp 608\u2013623","DOI":"10.1007\/978-3-031-25072-9_41"},{"key":"10869_CR31","doi-asserted-by":"crossref","unstructured":"Li Y, Quan R, Zhu L, Yang Y (2023) Efficient multimodal fusion via interactive prompting. In: Proceedings of CVPR, pp 2604\u20132613","DOI":"10.1109\/CVPR52729.2023.00256"},{"key":"10869_CR32","unstructured":"Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, Levy O, Lewis M, Zettlemoyer L, Stoyanov V (2019) RoBERTa: a robustly optimized BERT pretraining approach. arXiv preprint. arXiv:1907.11692"},{"key":"10869_CR33","doi-asserted-by":"publisher","first-page":"17","DOI":"10.1016\/j.patrec.2023.02.024","volume":"168","author":"Z Liu","year":"2023","unstructured":"Liu Z, Cheng Q, Song C, Cheng J (2023) Cross-scale cascade transformer for multimodal human action recognition. Pattern Recogn Lett 168:17\u201323","journal-title":"Pattern Recogn Lett"},{"key":"10869_CR34","doi-asserted-by":"crossref","unstructured":"Majumdar S, Ginsburg B (2020) MatchboxNet: 1D time-channel separable convolutional neural network architecture for speech commands recognition. In: Proceedings of INTERSPEECH, pp 3356\u20133360","DOI":"10.21437\/Interspeech.2020-1058"},{"key":"10869_CR35","doi-asserted-by":"crossref","unstructured":"Martin L, Muller B, Su\u00e1rez PJO, Dupont Y, Romary L, La\u00a0Clergerie \u00c9V, Seddah D, Sagot B (2020) CamemBERT: a Tasty French language model. In: Proceedings of ACL, pp 7203\u20137219","DOI":"10.18653\/v1\/2020.acl-main.645"},{"issue":"3","key":"10869_CR36","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3439726","volume":"54","author":"S Minaee","year":"2021","unstructured":"Minaee S, Kalchbrenner N, Cambria E, Nikzad N, Chenaghlu M, Gao J (2021) Deep learning-based text classification: a comprehensive review. ACM Comput Surv 54(3):1\u201340","journal-title":"ACM Comput Surv"},{"issue":"14","key":"10869_CR37","doi-asserted-by":"publisher","first-page":"4927","DOI":"10.3390\/s21144927","volume":"21","author":"YR Pandeya","year":"2021","unstructured":"Pandeya YR, Bhattarai B, Lee J (2021) Deep-learning-based multimodal emotion classification for music videos. Sensors 21(14):4927","journal-title":"Sensors"},{"issue":"5","key":"10869_CR38","doi-asserted-by":"publisher","first-page":"2381","DOI":"10.3390\/s23052381","volume":"23","author":"M Paw\u0142owski","year":"2023","unstructured":"Paw\u0142owski M, Wr\u00f3blewska A, Sysko-Roma\u0144czuk S (2023) Effective techniques for multimodal data fusion: a comparative analysis. Sensors 23(5):2381","journal-title":"Sensors"},{"key":"10869_CR39","doi-asserted-by":"crossref","unstructured":"Porjazovski D, Getman Y, Gr\u00f3sz T, Kurimo M (2023) Advancing audio emotion and intent recognition with large pre-trained models and Bayesian inference. In: Proceedings of ACMMM, pp 9477\u20139481","DOI":"10.1145\/3581783.3612848"},{"issue":"2","key":"10869_CR40","doi-asserted-by":"publisher","first-page":"206","DOI":"10.1109\/JSTSP.2019.2908700","volume":"13","author":"H Purwins","year":"2019","unstructured":"Purwins H, Li B, Virtanen T, Schl\u00fcter J, Chang S-Y, Sainath T (2019) Deep learning for audio signal processing. IEEE J Sel Top Signal Process 13(2):206\u2013219","journal-title":"IEEE J Sel Top Signal Process"},{"key":"10869_CR41","unstructured":"Radford A, Kim JW, Xu T, Brockman G, McLeavey C, Sutskever I (2022) Robust speech recognition via large-scale weak supervision. arXiv preprint. arXiv:2212.04356"},{"issue":"6","key":"10869_CR42","doi-asserted-by":"publisher","first-page":"96","DOI":"10.1109\/MSP.2017.2738401","volume":"34","author":"D Ramachandram","year":"2017","unstructured":"Ramachandram D, Taylor GW (2017) Deep multimodal learning: a survey on recent advances and trends. IEEE Signal Process Mag 34(6):96\u2013108","journal-title":"IEEE Signal Process Mag"},{"key":"10869_CR43","doi-asserted-by":"crossref","unstructured":"Ristea N-C, Ionescu RT (2020) Are you wearing a mask? Improving mask detection from speech using augmentation by cycle-consistent GANs. In: Proceedings of INTERSPEECH, pp 2102\u20132106","DOI":"10.21437\/Interspeech.2020-1329"},{"key":"10869_CR44","doi-asserted-by":"crossref","unstructured":"Ristea N-C, Ionescu RT (2023) Cascaded cross-modal transformer for request and complaint detection. In: Proceedings of ACMMM, pp 9467\u20139471","DOI":"10.1145\/3581783.3612846"},{"key":"10869_CR45","doi-asserted-by":"crossref","unstructured":"Ristea NC, Ionescu RT, Khan F (2022) SepTr: separable transformer for audio spectrogram processing. In: Proceedings of INTERSPEECH, pp 4103\u20134107","DOI":"10.21437\/Interspeech.2022-249"},{"key":"10869_CR46","doi-asserted-by":"crossref","unstructured":"Schuller BW, Batliner A, Amiriparian S, Barnhill A, Gerczuk M, Triantafyllopoulos A, Baird A, Tzirakis P, Gagne C, Cowen AS, Lackovic N, Caraty M-J, Montaci\u00e9 C (2023) The ACM multimedia 2023 computational paralinguistics challenge: emotion share & requests. In: Proceedings of ACMMM, pp 9635\u20139639","DOI":"10.1145\/3581783.3612835"},{"issue":"31","key":"10869_CR47","doi-asserted-by":"publisher","first-page":"22935","DOI":"10.1007\/s00521-022-06913-2","volume":"35","author":"A Sharma","year":"2023","unstructured":"Sharma A, Sharma K, Kumar A (2023) Real-time emotional health detection using fine-tuned transfer networks with multimodal fusion. Neural Comput Appl 35(31):22935\u201322948","journal-title":"Neural Comput Appl"},{"key":"10869_CR48","doi-asserted-by":"crossref","unstructured":"Shvetsova N, Chen B, Rouditchenko A, Thomas S, Kingsbury B, Feris RS, Harwath D, Glass J, Kuehne H (2022) Everything at once-multi-modal fusion transformer for video retrieval. In: Proceedings of CVPR, pp 20020\u201320029","DOI":"10.1109\/CVPR52688.2022.01939"},{"key":"10869_CR49","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.107316","volume":"229","author":"P Singh","year":"2021","unstructured":"Singh P, Srivastava R, Rana KPS, Kumar V (2021) A multimodal hierarchical approach to speech emotion recognition from audio and text. Knowl Based Syst 229:107316","journal-title":"Knowl Based Syst"},{"issue":"2","key":"10869_CR51","doi-asserted-by":"publisher","first-page":"569","DOI":"10.1093\/bib\/bbab569","volume":"23","author":"SR Stahlschmidt","year":"2022","unstructured":"Stahlschmidt SR, Ulfenborg B, Synnergren J (2022) Multimodal deep learning for biomedical data fusion: a review. Brief Bioinform 23(2):569","journal-title":"Brief Bioinform"},{"key":"10869_CR52","doi-asserted-by":"crossref","unstructured":"Sun Z, Sarma P, Sethares W, Liang Y (2020) Learning relationships between text, audio, and video via deep canonical correlation for multimodal language analysis. In: Proceedings of AAAI, vol 34, pp 8992\u20138999","DOI":"10.1609\/aaai.v34i05.6431"},{"key":"10869_CR53","doi-asserted-by":"crossref","unstructured":"Sun Y, Xu K, Liu C, Dou Y, Qian K (2023) Automatic audio augmentation for requests sub-challenge. In: Proceedings of ACMMM, pp 9482\u20139486","DOI":"10.1145\/3581783.3612849"},{"key":"10869_CR54","doi-asserted-by":"crossref","unstructured":"Sunder V, Thomas S, Kuo H-KJ, Ganhotra J, Kingsbury B, Fosler-Lussier E (2022) Towards end-to-end integration of dialog history for improved spoken language understanding. In: Proceedings of ICASSP. IEEE, pp 7497\u20137501","DOI":"10.1109\/ICASSP43922.2022.9747871"},{"key":"10869_CR55","doi-asserted-by":"crossref","unstructured":"Thomas S, Kuo H-KJ, Kingsbury B, Saon G (2022) Towards reducing the need for speech training data to build spoken language understanding systems. In: Proceedings of ICASSP. IEEE, pp 7932\u20137936","DOI":"10.1109\/ICASSP43922.2022.9747555"},{"key":"10869_CR56","doi-asserted-by":"crossref","unstructured":"Toto E, Tlachac ML, Rundensteiner EA (2021) AudiBERT: a deep transfer learning multimodal classification framework for depression screening. In: Proceedings of CIKM, pp 4145\u20134154","DOI":"10.1145\/3459637.3481895"},{"key":"10869_CR57","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. In: Proceedings of NIPS, pp 5998\u20136008"},{"issue":"4","key":"10869_CR58","first-page":"1","volume":"78","author":"C-X Wan","year":"2022","unstructured":"Wan C-X, Li B (2022) Financial causal sentence recognition based on BERT-CNN text classification. J Supercomput 78(4):1\u201325","journal-title":"J Supercomput"},{"key":"10869_CR59","unstructured":"Wang Y, Huang W, Sun F, Xu T, Rong Y, Huang J (2020) Deep multimodal fusion by channel exchanging. In: Proceedings of NeurIPS, vol 33, pp 4835\u20134845"},{"key":"10869_CR60","unstructured":"Warden P (2018) Speech commands: a dataset for limited-vocabulary speech recognition. arXiv preprint. arXiv:1804.03209"},{"key":"10869_CR61","unstructured":"Wu M, Nafziger J, Scodary A, Maas A (2020) HarperValleyBank: a domain-specific spoken dialog corpus. arXiv preprint. arXiv:2010.13929"},{"issue":"10","key":"10869_CR62","doi-asserted-by":"publisher","first-page":"12113","DOI":"10.1109\/TPAMI.2023.3275156","volume":"45","author":"P Xu","year":"2023","unstructured":"Xu P, Zhu X, Clifton DA (2023) Multimodal learning with transformers: a survey. IEEE Trans Pattern Anal Mach Intell 45(10):12113\u201312132","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"10869_CR63","doi-asserted-by":"crossref","unstructured":"Yang C-HH, Qi J, Chen SY-C, Tsao Y, Chen P-Y (2022) When BERT meets quantum temporal convolution learning for text classification in heterogeneous computing. In: Proceedings of ICASSP. IEEE, pp 8602\u20138606","DOI":"10.1109\/ICASSP43922.2022.9746412"},{"key":"10869_CR64","doi-asserted-by":"crossref","unstructured":"Yoon S, Byun S, Jung K (2018) Multimodal speech emotion recognition using audio and text. In: Proceedings of SLT workshop. IEEE, pp 112\u2013118","DOI":"10.1109\/SLT.2018.8639583"},{"issue":"6","key":"10869_CR65","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3414685.3417838","volume":"39","author":"Y Yoon","year":"2020","unstructured":"Yoon Y, Cha B, Lee J-H, Jang M, Lee J, Kim J, Lee G (2020) Speech gesture generation from the trimodal context of text, audio, and speaker identity. ACM Trans Graph 39(6):1\u201316","journal-title":"ACM Trans Graph"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-024-10869-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-024-10869-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-024-10869-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,5]],"date-time":"2024-09-05T05:05:19Z","timestamp":1725512719000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-024-10869-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,2]]},"references-count":64,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2024,9]]}},"alternative-id":["10869"],"URL":"https:\/\/doi.org\/10.1007\/s10462-024-10869-1","relation":{},"ISSN":["1573-7462"],"issn-type":[{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,2]]},"assertion":[{"value":"20 July 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 August 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no conflict of interest to declare that are relevant to the content of this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"The authors give their consent for publication.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}],"article-number":"225"}}