{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,6]],"date-time":"2026-02-06T03:39:57Z","timestamp":1770349197820,"version":"3.49.0"},"reference-count":84,"publisher":"American Association for the Advancement of Science (AAAS)","funder":[{"DOI":"10.13039\/501100020089","name":"Science and Technology Commission of Fengxian District, Shanghai Municipality","doi-asserted-by":"publisher","award":["22511105901"],"award-info":[{"award-number":["22511105901"]}],"id":[{"id":"10.13039\/501100020089","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["spj.science.org"],"crossmark-restriction":true},"short-container-title":["Intell Comput"],"published-print":{"date-parts":[[2024,1]]},"abstract":"<jats:p>Recent advances in self-supervised models have led to effective pretrained speech representations in downstream speech emotion recognition tasks. However, previous research has primarily focused on exploiting pretrained representations by simply adding a linear head on top of the pretrained model, while overlooking the design of the downstream network. In this paper, we propose a temporal shift module with pretrained representations to integrate channel-wise information without introducing additional parameters or floating-point operations per second. By incorporating the temporal shift module, we developed corresponding shift variants for 3 baseline building blocks: ShiftCNN, ShiftLSTM, and Shiftformer. Furthermore, we propose 2 technical strategies, placement and proportion of shift, to balance the trade-off between mingling and misalignment. Our family of temporal shift models outperforms state-of-the-art methods on the benchmark Interactive Emotional Dyadic Motion Capture dataset in fine-tuning and feature-extraction scenarios. In addition, through comprehensive experiments using wav2vec 2.0 and Hidden-Unit Bidirectional Encoder Representations from Transformers representations, we identified the behavior of the temporal shift module in downstream models, which may serve as an empirical guideline for future exploration of channel-wise shift and downstream network design.<\/jats:p>","DOI":"10.34133\/icomputing.0073","type":"journal-article","created":{"date-parts":[[2024,2,12]],"date-time":"2024-02-12T12:11:50Z","timestamp":1707739910000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark_01","source":"Crossref","is-referenced-by-count":9,"title":["Temporal Shift Module with Pretrained Representations for Speech Emotion Recognition"],"prefix":"10.34133","volume":"3","author":[{"given":"Siyuan","family":"Shen","sequence":"first","affiliation":[{"name":"Institute of AI for Education, \rEast China Normal University, Shanghai, China."},{"name":"School of Computer Science and Technology, \rEast China Normal University, Shanghai, China."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5289-5761","authenticated-orcid":false,"given":"Feng","family":"Liu","sequence":"additional","affiliation":[{"name":"Institute of AI for Education, \rEast China Normal University, Shanghai, China."},{"name":"School of Computer Science and Technology, \rEast China Normal University, Shanghai, China."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hanyang","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of AI for Education, \rEast China Normal University, Shanghai, China."},{"name":"School of Computer Science and Technology, \rEast China Normal University, Shanghai, China."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunlong","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Acoustics, \rUniversity of Chinese Academy of Sciences, Beijing, China."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4768-5946","authenticated-orcid":true,"given":"Aimin","family":"Zhou","sequence":"additional","affiliation":[{"name":"Institute of AI for Education, \rEast China Normal University, Shanghai, China."},{"name":"School of Computer Science and Technology, \rEast China Normal University, Shanghai, China."}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"221","published-online":{"date-parts":[[2024,2,21]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"crossref","unstructured":"Wu X Hu S Wu Z Liu X Meng H. Neural architecture search for speech emotion recognition. Paper presented at: ICASSP 2022\u20132022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2022 May 23\u201327; Singapore Singapore.","DOI":"10.1109\/ICASSP43922.2022.9746155"},{"issue":"8","key":"e_1_3_3_3_2","doi-asserted-by":"crossref","first-page":"2203","DOI":"10.1109\/TMM.2014.2360798","article-title":"Learning salient features for speech emotion recognition using convolutional neural networks","volume":"16","author":"Mao Q","year":"2014","unstructured":"Mao Q, Dong M, Huang Z, Zhan Y. Learning salient features for speech emotion recognition using convolutional neural networks. IEEE Trans Multimed. 2014;16(8):2203\u20132213.","journal-title":"IEEE Trans Multimed"},{"key":"e_1_3_3_4_2","doi-asserted-by":"crossref","unstructured":"Trigeorgis G Ringeval F Brueckner R Marchi E Nicolaou MA Schuller B Zafeiriou S. Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network. Paper presented at: 2016 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2016 Mar 20\u201325; Shanghai China.","DOI":"10.1109\/ICASSP.2016.7472669"},{"key":"e_1_3_3_5_2","doi-asserted-by":"crossref","unstructured":"Mirsamadi S Barsoum E Zhang C. Automatic speech emotion recognition using recurrent neural networks with local attention. Paper presented at: 2017 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2017 Mar 05\u201309; New Orleans LA USA.","DOI":"10.1109\/ICASSP.2017.7952552"},{"key":"e_1_3_3_6_2","first-page":"1517","article-title":"A database of German emotional speech","volume":"5","author":"Burkhardt F","year":"2005","unstructured":"Burkhardt F, Paeschke A, Rolfes M, Sendlmeier WF, Weiss B, et al. A database of German emotional speech. In: Interspeech. 2005;5:1517\u20131520.","journal-title":"In: Interspeech."},{"key":"e_1_3_3_7_2","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: Interactive emotional dyadic motion capture database","volume":"42","author":"Busso C","year":"2008","unstructured":"Busso C, Bulut M, Lee C-C, Kazemzadeh A, Mower E, Kim S, Chang JN, Lee S, Narayanan SS. IEMOCAP: Interactive emotional dyadic motion capture database. Lang Resour Eval. 2008;42:335\u2013359.","journal-title":"Lang Resour Eval"},{"issue":"5","key":"e_1_3_3_8_2","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pone.0196391","article-title":"The Ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in north American English","volume":"13","author":"Livingstone SR","year":"2018","unstructured":"Livingstone SR, Russo FA. The Ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in north American English. PLoS One. 2018;13(5): Article e0196391.","journal-title":"PLoS One"},{"issue":"2","key":"e_1_3_3_9_2","doi-asserted-by":"crossref","first-page":"1634","DOI":"10.1109\/TAFFC.2021.3114365","article-title":"Survey of deep representation learning for speech emotion recognition","volume":"14","author":"Latif S","year":"2021","unstructured":"Latif S, Rana R, Khalifa S, Jurdak R, Qadir J, Schuller B. Survey of deep representation learning for speech emotion recognition. IEEE Trans Affect Comput. 2021;14(2):1634\u20131654.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_10_2","first-page":"12449","article-title":"wav2vec 2.0: A framework for self-supervised learning of speech representations","volume":"33","author":"Baevski A","year":"2020","unstructured":"Baevski A, Zhou Y, Mohamed A, Auli M. wav2vec 2.0: A framework for self-supervised learning of speech representations. Adv Neural Inf Proces Syst. 2020;33:12449\u201312460.","journal-title":"Adv Neural Inf Proces Syst"},{"key":"e_1_3_3_11_2","doi-asserted-by":"crossref","first-page":"3451","DOI":"10.1109\/TASLP.2021.3122291","article-title":"Hubert: Self-supervised speech representation learning by masked prediction of hidden units","volume":"29","author":"Hsu WN","year":"2021","unstructured":"Hsu WN, Bolte B, Tsai YHH, Lakhotia K, Salakhutdinov R, Mohamed A. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE\/ACM Trans Audio Speech Lang Process. 2021;29:3451\u20133460.","journal-title":"IEEE\/ACM Trans Audio Speech Lang Process"},{"key":"e_1_3_3_12_2","unstructured":"van den Oord A Li Y Vinyals O. Representation learning with contrastive predictive coding. ArXiv. 2018. https:\/\/doi.org\/10.48550\/arXiv.1807.03748"},{"key":"e_1_3_3_13_2","unstructured":"Devlin J Chang MW Lee K Toutanova K. Bert: Pre-training of deep bidirectional transformers for language understanding. ArXiv. 2018. https:\/\/doi.org\/10.48550\/arXiv.1810.04805"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.50"},{"issue":"6","key":"e_1_3_3_15_2","doi-asserted-by":"crossref","first-page":"1505","DOI":"10.1109\/JSTSP.2022.3188113","article-title":"WavLM: Large-scale self-supervised pre-training for full stack speech processing","volume":"564","author":"Chen S","year":"2022","unstructured":"Chen S, Wang C, Chen Z, Wu Y, Liu S, Chen Z, Li J, Kanda N, Yoshioka T, Xiao X, et al. WavLM: Large-scale self-supervised pre-training for full stack speech processing. IEEE J Sel Top Signal Process. 2022;564(6):1505\u20131518.","journal-title":"IEEE J Sel Top Signal Process"},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","first-page":"1194","DOI":"10.21437\/Interspeech.2021-1775","article-title":"SUPERB: Speech processing universal performance benchmark","author":"Yang S-W","year":"2021","unstructured":"Yang S-W, Chi P-H, Chuang Y-S, Lai C-IJ, Lakhotia K, Lin YY, Liu AT, Shi J, Chang X, Lin G-T, et al. SUPERB: Speech processing universal performance benchmark. Proc. Interspeech 2021. 2021;1194\u20131198.","journal-title":"Proc. Interspeech 2021"},{"key":"e_1_3_3_17_2","doi-asserted-by":"crossref","unstructured":"Boigne J Liyanage B Ostrem T. Recognizing more emotions with less data using self-supervised transfer learning. ArXiv. 2020. https:\/\/doi.org\/10.48550\/arXiv.2011.05585.","DOI":"10.20944\/preprints202008.0645.v1"},{"key":"e_1_3_3_18_2","doi-asserted-by":"crossref","first-page":"3370","DOI":"10.21437\/Interspeech.2021-1840","article-title":"Temporal context in speech emotion recognition","author":"Xia Y","year":"2021","unstructured":"Xia Y, Chen LW, Rudnicky A, Stern RM. Temporal context in speech emotion recognition. Proc Interspeech 2021. 2021;3370\u20133374.","journal-title":"Proc Interspeech 2021"},{"key":"e_1_3_3_19_2","doi-asserted-by":"crossref","first-page":"3400","DOI":"10.21437\/Interspeech.2021-703","article-title":"Emotion recognition from speech using wav2vec 2.0 embeddings","author":"Pepino L","year":"2021","unstructured":"Pepino L, Riera P, Ferrer L. Emotion recognition from speech using wav2vec 2.0 embeddings. Proc Interspeech 2021. 2021;3400\u20133404.","journal-title":"Proc Interspeech 2021"},{"key":"e_1_3_3_20_2","doi-asserted-by":"crossref","unstructured":"Chen L-W Rudnicky A. Exploring Wav2vec 2.0 fine tuning for improved speech emotion recognition. Paper presented at: ICASSP 2023\u20132023 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2023 Jun 04\u201310; Rhodes Island Greece.","DOI":"10.1109\/ICASSP49357.2023.10095036"},{"key":"e_1_3_3_21_2","doi-asserted-by":"crossref","unstructured":"Sharma M. Multi-lingual multi-task speech emotion recognition using wav2vec 2.0. Paper presented at: ICASSP 2022\u20132022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2022 May 23\u201327; Singapore Singapore.","DOI":"10.1109\/ICASSP43922.2022.9747417"},{"key":"e_1_3_3_22_2","doi-asserted-by":"crossref","unstructured":"Huang B Carley KM. Syntax-aware aspect level sentiment classification with graph attention networks. Paper presented at: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP); 2019 Nov 3\u20137; Hongkong China.","DOI":"10.18653\/v1\/D19-1549"},{"issue":"11","key":"e_1_3_3_23_2","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun Y","year":"1998","unstructured":"LeCun Y, Bottou L, Bengio Y, Haffner P. Gradient-based learning applied to document recognition. Proc IEEE. 1998;86(11):2278\u20132324.","journal-title":"Proc IEEE"},{"key":"e_1_3_3_24_2","doi-asserted-by":"crossref","unstructured":"Rumelhart DE Hinton GE Williams RJ. Learning internal representations by error propagation. San Diego (CA): California University Institute for Cognitive Science; 1985.","DOI":"10.21236\/ADA164453"},{"key":"e_1_3_3_25_2","unstructured":"Vaswani A Shazeer N Parmar N Uszkoreit J Jones L Gomez AN Kaiser \u0141 Polosukhin I. Attention is all you need. Paper presented at: 31st Conference on Neural Information Processing Systems (NIPS); 2017 Dec 4\u20139; Long Beach CA USA."},{"key":"e_1_3_3_26_2","doi-asserted-by":"crossref","unstructured":"Liu Z Mao H Wu C-Y Feichtenhofer C Darrell T Xie S. A convnet for the 2020s. Paper presented at: 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 Jun 18\u201324; New Orleans LA USA.","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"e_1_3_3_27_2","doi-asserted-by":"crossref","unstructured":"Graves A Fernandez S Gomez F Schmidhuber J. Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. In: Proceedings of the 23rd international conference on machine learning. Association for Computing Machinery; 2006. p. 369\u2013376.","DOI":"10.1145\/1143844.1143891"},{"key":"e_1_3_3_28_2","unstructured":"Sequence transduction with recurrent neural networks. ArXiv. 2012. https:\/\/doi.org\/10.48550\/arXiv.1211.3711"},{"key":"e_1_3_3_29_2","doi-asserted-by":"crossref","unstructured":"Graves A Mohamed A-R Hinton G. Speech recognition with deep recurrent neural networks. Paper presented at: 2013 IEEE international conference on acoustics speech and signal processing; 2013 May 26\u201331; Vancouver BC Canada.","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"e_1_3_3_30_2","doi-asserted-by":"crossref","unstructured":"Shen S Liu F Zhou A. Mingling or misalignment? Temporal shift for speech emotion recognition with pre-trained representations. Paper presented at: ICASSP 2023-2023 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2023 Jun 04\u201310; Rhodes Island Greece.","DOI":"10.1109\/ICASSP49357.2023.10095193"},{"issue":"6","key":"e_1_3_3_31_2","doi-asserted-by":"crossref","first-page":"1576","DOI":"10.1109\/TMM.2017.2766843","article-title":"Speech emotion recognition using deep convolutional neural network and discriminant temporal pyramid matching","volume":"20","author":"Zhang S","year":"2017","unstructured":"Zhang S, Zhang S, Huang T, Gao W. Speech emotion recognition using deep convolutional neural network and discriminant temporal pyramid matching. IEEE Trans Multimed. 2017;20(6):1576\u20131590.","journal-title":"IEEE Trans Multimed"},{"key":"e_1_3_3_32_2","doi-asserted-by":"crossref","unstructured":"Zheng WQ Yu JS Zou YX. An experimental study of speech emotion recognition based on deep convolutional neural networks. Paper presented at: 2015 International Conference on Affective Computing and Intelligent Interaction (ACII); 2015 Sep 21\u201324; Xi\u2019an China.","DOI":"10.1109\/ACII.2015.7344669"},{"issue":"4","key":"e_1_3_3_33_2","doi-asserted-by":"crossref","first-page":"1813","DOI":"10.1109\/TCSS.2022.3199119","article-title":"OPO-FCM: A computational affection based OCC-PAD-OCEAN federation cognitive modeling approach","volume":"10","author":"Liu F","year":"2022","unstructured":"Liu F, Wang H-Y, Shen S-Y, Jia X, Hu J-Y, Zhang J-H, Wang X-Y, Lei Y, Zhou A-M, Qi J-Y, et al. OPO-FCM: A computational affection based OCC-PAD-OCEAN federation cognitive modeling approach. IEEE Trans Comput Soc Syst. 2022;10(4):1813\u20131825.","journal-title":"IEEE Trans Comput Soc Syst"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.bspc.2018.08.035"},{"key":"e_1_3_3_35_2","doi-asserted-by":"crossref","unstructured":"Wang H Li B Wu S Shen S Liu F Ding S Zhou A. Rethinking the learning paradigm for dynamic facial expression recognition. Paper presented at: 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2023 Jun 17\u201324; Vancouver BC Canada.","DOI":"10.1109\/CVPR52729.2023.01722"},{"key":"e_1_3_3_36_2","first-page":"2803","article-title":"Improved end-to-end speech emotion recognition using self attention mechanism and multitask learning","author":"Li Y","year":"2019","unstructured":"Li Y, Zhao T, Kawahara T. Improved end-to-end speech emotion recognition using self attention mechanism and multitask learning. Interspeech. 2019;2803\u20132807.","journal-title":"Interspeech"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2022.3207050"},{"key":"e_1_3_3_38_2","first-page":"161","article-title":"Learning problem-agnostic speech representations from multiple self-supervised tasks","volume":"2019","author":"Pascual S","year":"2019","unstructured":"Pascual S, Ravanelli M, Serra J, Bonafonte A, Bengio Y. Learning problem-agnostic speech representations from multiple self-supervised tasks. Proc Interspeech. 2019;2019:161\u2013165.","journal-title":"Proc Interspeech"},{"key":"e_1_3_3_39_2","doi-asserted-by":"crossref","unstructured":"Ravanelli M Zhong J Pascual S Swietojanski P Monteiro J Trmal J Bengio Y. Multi-task self-supervised learning for robust speech recognition. Paper presented at: ICASSP 2020-2020 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2020 May 04\u201308; Barcelona Spain.","DOI":"10.1109\/ICASSP40776.2020.9053569"},{"issue":"2","key":"e_1_3_3_40_2","doi-asserted-by":"crossref","first-page":"190","DOI":"10.1109\/TAFFC.2015.2457417","article-title":"The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing","volume":"7","author":"Eyben F","year":"2015","unstructured":"Eyben F, Scherer KR, Schuller BW, Sundberg J, Andre E, Busso C, Devillers LY, Epps J, Laukka P, Narayanan SS, et al. The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing. IEEE Trans Affect Comput. 2015;7(2):190\u2013202.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_41_2","doi-asserted-by":"crossref","unstructured":"Li Y Bell P Lai C. Fusing ASR outputs in joint training for speech emotion recognition. Paper presented at: ICASSP 2022-2022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2022 May 23\u201327; Singapore Singapore.","DOI":"10.1109\/ICASSP43922.2022.9746289"},{"key":"e_1_3_3_42_2","doi-asserted-by":"crossref","unstructured":"Gat I Aronowitz H Zhu W Morais E Hoory R. Speaker normalization for self-supervised speech emotion recognition. Paper presented at: ICASSP 2022-2022 IEEE International Conference on Acoustics Speech and Signal processing (ICASSP); 2022 May 23\u201327; Singapore Singapore.","DOI":"10.1109\/ICASSP43922.2022.9747460"},{"key":"e_1_3_3_43_2","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1016\/j.specom.2019.12.001","article-title":"Speech emotion recognition: Emotional models, databases, features, preprocessing methods, supporting modalities, and classifiers","volume":"116","author":"Ak\u00e7ay MB","year":"2020","unstructured":"Ak\u00e7ay MB, O\u011fuz K. Speech emotion recognition: Emotional models, databases, features, preprocessing methods, supporting modalities, and classifiers. Speech Common. 2020;116:56\u201376.","journal-title":"Speech Common"},{"key":"e_1_3_3_44_2","doi-asserted-by":"crossref","first-page":"2489","DOI":"10.1109\/TASLP.2020.3016487","article-title":"On cross-corpus generalization of deep learning based speech enhancement","volume":"28","author":"Pandey A","year":"2020","unstructured":"Pandey A, Wang D. On cross-corpus generalization of deep learning based speech enhancement. IEEE\/ACM Trans Audio Speech Lang Process. 2020;28:2489\u20132499.","journal-title":"IEEE\/ACM Trans Audio Speech Lang Process"},{"key":"e_1_3_3_45_2","doi-asserted-by":"crossref","first-page":"328","DOI":"10.1109\/29.21701","article-title":"Phoneme recognition using time-delay neural networks","volume":"37","author":"Waibel A","year":"1989","unstructured":"Waibel A, Hanazawa T, Hinton G, Shikano K, Lang KJ. Phoneme recognition using time-delay neural networks. IEEE Trans Acoust Speech Signal Process. 1989;37:328\u2013339.","journal-title":"IEEE Trans Acoust Speech Signal Process"},{"key":"e_1_3_3_46_2","doi-asserted-by":"crossref","unstructured":"Xia S Fourer D Audin L Rouas JL Shochi T. Speech emotion recognition using time frequency random circular shift and deep neural networks. Paper presented at: Speech Prosody 2022; 2022 May 23; Lisbon Portugal.","DOI":"10.21437\/SpeechProsody.2022-119"},{"key":"e_1_3_3_47_2","first-page":"3859","article-title":"CyclicAugment: Speech data random augmentation with cosine annealing scheduler for automatic speech recognition","author":"Wang Z","unstructured":"Wang Z, Hou F, Qiu Y, Ma Z, Singh S, Wang R. CyclicAugment: Speech data random augmentation with cosine annealing scheduler for automatic speech recognition. Proc. Interspeech 2022. 3859\u20133863.","journal-title":"Proc. Interspeech 2022"},{"key":"e_1_3_3_48_2","doi-asserted-by":"crossref","unstructured":"Wu B Wan A Yue X Jin P Zhao S Golmant N Gholaminejad A Gonzalez J Keutzer K. Shift: A zero flop zero parameter alternative to spatial convolutions. Paper presented at: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2018 Jun 18\u201323; Salt Lake City UT USA.","DOI":"10.1109\/CVPR.2018.00951"},{"key":"e_1_3_3_49_2","doi-asserted-by":"crossref","unstructured":"Wang G Zhao Y Tang C Luo C Zeng W. When shift operation meets vision transformer: An extremely simple alternative to attention mechanism. In: Proceedings of the AAAI conference on artificial intelligence. EAAI 2022 Virtual Event: AAAI Press; 2022. p. 2423\u20132430.","DOI":"10.1609\/aaai.v36i2.20142"},{"key":"e_1_3_3_50_2","doi-asserted-by":"crossref","unstructured":"Lin J Gan C Han S. Tsm: Temporal shift module for efficient video understanding. Paper presented at: Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV); 2019 Oct 27\u201302 Nov; Seoul Korea (South).","DOI":"10.1109\/ICCV.2019.00718"},{"key":"e_1_3_3_51_2","unstructured":"Howard AG Zhu M Chen B Kalenichenko D Wang W Weyand T Andreetto M Adam H. Mobilenets: Efficient convolutional neural networks for mobile vision applications. ArXiv. 2017. https:\/\/doi.org\/10.48550\/arXiv.1704.04861."},{"key":"e_1_3_3_52_2","doi-asserted-by":"crossref","unstructured":"Chollet F. Xception: Deep learning with depthwise separable convolutions. Paper presented at: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017 Jul 21\u201326; Honolulu HI USA.","DOI":"10.1109\/CVPR.2017.195"},{"key":"e_1_3_3_53_2","doi-asserted-by":"crossref","unstructured":"Kriman S Beliaev S Ginsburg B Huang J Kuchaiev O Lavrukhin V Leary R Li J Zhang Y. Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions. Paper presented at: ICASSP 2020-2020 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2020 May 04\u201308; Barcelona Spain.","DOI":"10.1109\/ICASSP40776.2020.9053889"},{"key":"e_1_3_3_54_2","unstructured":"Chen S Xie E Ge C Chen R Liang D Luo P. CycleMLP: A MLP-like architecture for dense prediction. ArXiv. 2022. https:\/\/doi.org\/10.48550\/arXiv.2107.10224"},{"key":"e_1_3_3_55_2","doi-asserted-by":"crossref","first-page":"5036","DOI":"10.21437\/Interspeech.2020-3015","article-title":"Conformer: Convolution-augmented transformer for speech recognition","author":"Gulati A","year":"2020","unstructured":"Gulati A, Qin J, Chiu CC, Parmar N, Zhang Y, Yu J, Han W, Wang S, Zhang Z, Wu Y, et al. Conformer: Convolution-augmented transformer for speech recognition. Proc Interspeech 2020. 2020;5036\u20135040.","journal-title":"Proc Interspeech 2020"},{"key":"e_1_3_3_56_2","doi-asserted-by":"crossref","unstructured":"Yu W Luo M Zhou P Si C Zhou Y Wang X Feng J Yan S. Metaformer is actually what you need for vision. Paper presented at: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 Jun 18\u201324; New Orleans LA USA.","DOI":"10.1109\/CVPR52688.2022.01055"},{"key":"e_1_3_3_57_2","first-page":"24261","article-title":"Mlp-mixer: An all-mlp architecture for vision","volume":"34","author":"Tolstikhin IO","year":"2021","unstructured":"Tolstikhin IO, Houlsby N, Kolesnikov A, Beyer L, Zhai X, Unterthiner T, Yung J, Steiner A, Keysers D, Uszkoreit J, et al. Mlp-mixer: An all-mlp architecture for vision. Adv Neural Inf Proces Syst. 2021;34:24261\u201324272.","journal-title":"Adv Neural Inf Proces Syst"},{"key":"e_1_3_3_58_2","doi-asserted-by":"crossref","unstructured":"Sandler M Howard A Zhu M Zhmoginov A Chen L-C. Mobilenetv2: Inverted residuals and linear bottlenecks. Paper presented at: Proceedings of the IEEE conference on computer vision and pattern recognition; 2018 Jun 18\u201323; Salt Lake City UT USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_3_59_2","doi-asserted-by":"crossref","first-page":"3610","DOI":"10.21437\/Interspeech.2020-2059","article-title":"ContextNet: Improving convolutional neural networks for automatic speech recognition with global context","author":"Han W","year":"2020","unstructured":"Han W, Zhang Z, Zhang Y, Yu J, Chiu C-C, Qin J, Gulati A, Pang R, Wu Y. ContextNet: Improving convolutional neural networks for automatic speech recognition with global context. Proc Interspeech 2020. 2020;3610\u20133614.","journal-title":"Proc Interspeech 2020"},{"key":"e_1_3_3_60_2","doi-asserted-by":"crossref","first-page":"3830","DOI":"10.21437\/Interspeech.2020-2650","article-title":"ECAPA-TDNN: Emphasized Channel attention, propagation and aggregation in TDNN based speaker verification","author":"Desplanques B","year":"2020","unstructured":"Desplanques B, Thienpondt J, Demuynck K. ECAPA-TDNN: Emphasized Channel attention, propagation and aggregation in TDNN based speaker verification. Proc Interspeech 2020. 2020;3830\u20133834.","journal-title":"Proc Interspeech 2020"},{"key":"e_1_3_3_61_2","unstructured":"Ioffe S Szegedy C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: International conference on machine learning. Lille (France): PMLR; 2015. p. 448\u2013456."},{"key":"e_1_3_3_62_2","unstructured":"Ba JL Kiros JR Hinton GE. Layer normalization. arXiv. 2016. https:\/\/doi.org\/10.48550\/arXiv.1607.06450."},{"key":"e_1_3_3_63_2","doi-asserted-by":"crossref","unstructured":"Li N Liu S Liu Y Zhao S Liu M. Neural speech synthesis with transformer network. In: Proceedings of the AAAI conference on artificial intelligence. Honolulu (HI): AAAI Press; 2019. p. 6706\u20136713.","DOI":"10.1609\/aaai.v33i01.33016706"},{"key":"e_1_3_3_64_2","doi-asserted-by":"crossref","unstructured":"Dai Z Yang Z Yang Y Carbonell JG Le Q Salakhutdinov R. Transformer-XL: Attentive language models beyond a fixed-length context. In: Proceedings of the 57th annual meeting of the association for computational linguistics. Lille (France): JMLR.org; 2019. p. 2978\u20132988.","DOI":"10.18653\/v1\/P19-1285"},{"key":"e_1_3_3_65_2","unstructured":"Xiong R Yang Y He D Zheng K Zheng S Xing C Zhang H Lan Y Wang L Liu T-Y. On layer normalization in the transformer architecture. In: International conference on machine learning. PMLR; 2020. p. 10524\u201310533."},{"key":"e_1_3_3_66_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_3_67_2","doi-asserted-by":"crossref","unstructured":"Cho K van Merrienboer B Bahdanau D Bengio Y. On the properties of neural machine translation: Encoder\u2013decoder approaches. In: Proceedings of SSST-8 eighth workshop on syntax semantics and structure in statistical translation. Doha (Qatar): Association for Computational Linguistics; 2014. p. 103\u2013111.","DOI":"10.3115\/v1\/W14-4012"},{"key":"e_1_3_3_68_2","unstructured":"Zhang J Jia H. Design of speech corpus for mandarin text to speech. In: The blizzard challenge 2008 workshop. NLPR: 2008."},{"key":"e_1_3_3_69_2","unstructured":"Jackson P Haq S. Surrey audio-visual expressed emotion (savee) database. Guildford (UK): University of Surrey; 2014."},{"key":"e_1_3_3_70_2","doi-asserted-by":"crossref","unstructured":"Font F Roma G Serra X. Freesound technical demo. In: Proceedings of the 21st ACM international conference on multimedia. New York (NY): Association for Computing Machinery; 2013. p. 411\u2013412.","DOI":"10.1145\/2502081.2502245"},{"key":"e_1_3_3_71_2","doi-asserted-by":"crossref","unstructured":"Prasad A Jyothi P Velmurugan R. An investigation of end-to-end models for robust speech recognition. Paper presented at: ICASSP 2021\u20132021 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2021 Jun 06\u201311; Toronto ON Canada.","DOI":"10.1109\/ICASSP39728.2021.9414027"},{"key":"e_1_3_3_72_2","doi-asserted-by":"crossref","unstructured":"Zhu Q-S Zhang J Zhang Z-Q Wu M-H Fang X Dai L-R. A noise-robust self-supervised pre-training model based speech representation learning for automatic speech recognition. Paper presented at: ICASSP 2022\u20132022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2022 May 23\u201327; Singapore Singapore.","DOI":"10.1109\/ICASSP43922.2022.9747379"},{"issue":"12","key":"e_1_3_3_73_2","doi-asserted-by":"crossref","first-page":"2935","DOI":"10.1109\/TPAMI.2017.2773081","article-title":"Learning without forgetting","volume":"40","author":"Li Z","year":"2017","unstructured":"Li Z, Hoiem D. Learning without forgetting. IEEE Trans Pattern Anal Mach Intell. 2017;40(12):2935\u20132947.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"e_1_3_3_74_2","unstructured":"Kingma D Ba J. Adam: A method for stochastic optimization. Paper presented at: International Conference on Learning Representations (ICLR). 2015 May 7\u20139; San Diego CA USA."},{"key":"e_1_3_3_75_2","unstructured":"Loshchilov I Hutter F. Decoupled weight decay regularization. Paper presented at: International Conference on Learning Representations. 2019 May 6\u20139."},{"key":"e_1_3_3_76_2","doi-asserted-by":"crossref","first-page":"2613","DOI":"10.21437\/Interspeech.2019-2680","article-title":"SpecAugment: A simple data augmentation method for automatic speech recognition","volume":"2019","author":"Park DS","year":"2019","unstructured":"Park DS, Chan W, Zhang Y, Chiu C-C, Zoph B, Cubuk ED, Le QV. SpecAugment: A simple data augmentation method for automatic speech recognition. Proc Interspeech 2019. 2019;2019:2613\u20132617.","journal-title":"Proc Interspeech 2019"},{"key":"e_1_3_3_77_2","first-page":"550","article-title":"Belongie S. Residual networks behave like ensembles of relatively shallow networks","volume":"29","author":"Veit A","year":"2016","unstructured":"Veit A, Wilber MJ. Belongie S. Residual networks behave like ensembles of relatively shallow networks. Adv Neural Inf Process. 2016;29:550\u2013558.","journal-title":"Adv Neural Inf Process"},{"key":"e_1_3_3_78_2","doi-asserted-by":"crossref","unstructured":"Ding X Zhang X Han J Ding G. Scaling up your kernels to 31x31: Revisiting large kernel design in cnns. Paper presented at: 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 Jun 18\u201324; New Orleans LA USA.","DOI":"10.1109\/CVPR52688.2022.01166"},{"key":"e_1_3_3_79_2","doi-asserted-by":"crossref","unstructured":"He K Zhang X Ren S Sun J. Deep residual learning for image recognition. Paper presented at: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016 Jun 27\u201330; Las Vegas NV USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_80_2","doi-asserted-by":"crossref","unstructured":"Rajamani ST Rajamani KT Mallol-Ragolta A Liu S Schuller B. A novel attention-based gated recurrent unit and its efficacy in speech emotion recognition. Paper presented at: ICASSP 2021-2021 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2021 Jun 06\u201311; Toronto ON Canada.","DOI":"10.1109\/ICASSP39728.2021.9414489"},{"key":"e_1_3_3_81_2","doi-asserted-by":"crossref","unstructured":"Zou H Si Y Chen C Rajan D Chng ES. Speech emotion recognition with co-attention based multi-level acoustic information. Paper presented at: ICASSP 2022-2022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2022 May 23\u201327; Singapore Singapore.","DOI":"10.1109\/ICASSP43922.2022.9747095"},{"key":"e_1_3_3_82_2","doi-asserted-by":"crossref","unstructured":"Hu J Shen L Sun G. Squeeze-and-excitation networks. 2018 IEEE\/CVF conference on computer vision and pattern recognition; 2018 Jun 18\u201323; Salt Lake City UT USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_3_83_2","unstructured":"Ravanelli M Parcollet T Plantinga P Rouhe A Cornell S Lugosch L Subakan C Dawalatabad N Heba A Zhong J et\u00a0al. SpeechBrain: A general-purpose speech toolkit. ArXiv. 2021. https:\/\/doi.org\/10.48550\/arXiv.2106.04624"},{"key":"e_1_3_3_84_2","doi-asserted-by":"crossref","unstructured":"Lu C Zong Y Zheng W Li Y Tang C Schuller BW. Domain invariant feature learning for speaker-independent speech emotion recognition. In: IEEE\/ACM transactions on audio speech and language processing. IEEE; 2022. p. 2217\u20132230.","DOI":"10.1109\/TASLP.2022.3178232"},{"key":"e_1_3_3_85_2","doi-asserted-by":"crossref","unstructured":"Hershey S Chaudhuri S Ellis DPW Gemmeke JF Jansen A Moore RC Plakal M Platt D Saurous RA Seybold B et\u00a0al. CNN architectures for large-scale audio classification. Paper presented at: 2017 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2017 Mar 05\u201309; New Orleans LA USA.","DOI":"10.1109\/ICASSP.2017.7952132"}],"container-title":["Intelligent Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/spj.science.org\/doi\/pdf\/10.34133\/icomputing.0073","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,2,21]],"date-time":"2024-02-21T10:53:52Z","timestamp":1708512832000},"score":1,"resource":{"primary":{"URL":"https:\/\/spj.science.org\/doi\/10.34133\/icomputing.0073"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1]]},"references-count":84,"alternative-id":["10.34133\/icomputing.0073"],"URL":"https:\/\/doi.org\/10.34133\/icomputing.0073","relation":{},"ISSN":["2771-5892"],"issn-type":[{"value":"2771-5892","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1]]},"assertion":[{"value":"2023-07-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"0073"}}