{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T00:12:09Z","timestamp":1782951129597,"version":"3.54.5"},"reference-count":63,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2022,7,21]],"date-time":"2022-07-21T00:00:00Z","timestamp":1658361600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"2019, Digital Transformation in China and Germany: Strategies, Structures and Solutions for Ageing Societies","award":["GZ 1570"],"award-info":[{"award-number":["GZ 1570"]}]},{"name":"2019, Digital Transformation in China and Germany: Strategies, Structures and Solutions for Ageing Societies","award":["20dz2260300"],"award-info":[{"award-number":["20dz2260300"]}]},{"name":"Research Project of Shanghai Science and Technology Commission","award":["GZ 1570"],"award-info":[{"award-number":["GZ 1570"]}]},{"name":"Research Project of Shanghai Science and Technology Commission","award":["20dz2260300"],"award-info":[{"award-number":["20dz2260300"]}]},{"name":"The Fundamental Research Funds for the Central Universities","award":["GZ 1570"],"award-info":[{"award-number":["GZ 1570"]}]},{"name":"The Fundamental Research Funds for the Central Universities","award":["20dz2260300"],"award-info":[{"award-number":["20dz2260300"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Semantic-rich speech emotion recognition has a high degree of popularity in a range of areas. Speech emotion recognition aims to recognize human emotional states from utterances containing both acoustic and linguistic information. Since both textual and audio patterns play essential roles in speech emotion recognition (SER) tasks, various works have proposed novel modality fusing methods to exploit text and audio signals effectively. However, most of the high performance of existing models is dependent on a great number of learnable parameters, and they can only work well on data with fixed length. Therefore, minimizing computational overhead and improving generalization to unseen data with various lengths while maintaining a certain level of recognition accuracy is an urgent application problem. In this paper, we propose LGCCT, a light gated and crossed complementation transformer for multimodal speech emotion recognition. First, our model is capable of fusing modality information efficiently. Specifically, the acoustic features are extracted by CNN-BiLSTM while the textual features are extracted by BiLSTM. The modality-fused representation is then generated by the cross-attention module. We apply the gate-control mechanism to achieve the balanced integration of the original modality representation and the modality-fused representation. Second, the degree of attention focus can be considered, as the uncertainty and the entropy of the same token should converge to the same value independent of the length. To improve the generalization of the model to various testing-sequence lengths, we adopt the length-scaled dot product to calculate the attention score, which can be interpreted from a theoretical view of entropy. The operation of the length-scaled dot product is cheap but effective. Experiments are conducted on the benchmark dataset CMU-MOSEI. Compared to the baseline models, our model achieves an 81.0% F1 score with only 0.432 M parameters, showing an improvement in the balance between performance and the number of parameters. Moreover, the ablation study signifies the effectiveness of our model and its scalability to various input-sequence lengths, wherein the relative improvement is almost 20% of the baseline without a length-scaled dot product.<\/jats:p>","DOI":"10.3390\/e24071010","type":"journal-article","created":{"date-parts":[[2022,7,21]],"date-time":"2022-07-21T22:38:50Z","timestamp":1658443130000},"page":"1010","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["LGCCT: A Light Gated and Crossed Complementation Transformer for Multimodal Speech Emotion Recognition"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5289-5761","authenticated-orcid":false,"given":"Feng","family":"Liu","sequence":"first","affiliation":[{"name":"Institute of AI for Education, East China Normal University, Shanghai 200062, China"},{"name":"School of Computer Science and Technology, East China Normal University, Shanghai 200062, China"},{"name":"School of Computer Science, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Si-Yuan","family":"Shen","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, East China Normal University, Shanghai 200062, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zi-Wang","family":"Fu","sequence":"additional","affiliation":[{"name":"School of Computer Science, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Han-Yang","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, East China Normal University, Shanghai 200062, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ai-Min","family":"Zhou","sequence":"additional","affiliation":[{"name":"Institute of AI for Education, East China Normal University, Shanghai 200062, China"},{"name":"School of Computer Science and Technology, East China Normal University, Shanghai 200062, China"},{"name":"Shanghai Key Laboratory of Mental Health and Psychological Crisis Intervention, School of Psychology and Cognitive Science, East China Normal University, Shanghai 200062, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jia-Yin","family":"Qi","sequence":"additional","affiliation":[{"name":"Institute of Artificial Intelligence and Change Management, Shanghai University of International Business and Economics, Shanghai 200062, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,7,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1109\/MMUL.2021.3071495","article-title":"Implicit Emotion Relationship Mining Based on Optimal and Majority Synthesis from Multimodal Data Prediction","volume":"28","author":"Wang","year":"2021","journal-title":"IEEE MultiMedia"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Card, S.K., Moran, T.P., and Newell, A. (2018). The Psychology of Human-Computer Interaction, CRC Press.","DOI":"10.1201\/9780203736166"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Lugovi\u0107, S., Dun\u0111er, I., and Horvat, M. (Jun, January 30). Techniques and applications of emotion recognition in speech. Proceedings of the 2016 39th international convention on information and communication technology, electronics and microelectronics (mipro), Opatija, Croatia.","DOI":"10.1109\/MIPRO.2016.7522336"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1109\/79.911197","article-title":"Emotion recognition in human-computer interaction","volume":"18","author":"Cowie","year":"2001","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Tzirakis, P., Zhang, J., and Schuller, B.W. (2018, January 15\u201320). End-to-end speech emotion recognition using deep neural networks. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8462677"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Mirsamadi, S., Barsoum, E., and Zhang, C. (2017, January 5\u20139). Automatic speech emotion recognition using recurrent neural networks with local attention. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952552"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Satt, A., Rozenberg, S., and Hoory, R. (2017, January 20\u201324). Efficient Emotion Recognition from Speech Using Deep Learning on Spectrograms. Proceedings of the Interspeech, Stockholm, Sweden.","DOI":"10.21437\/Interspeech.2017-200"},{"key":"ref_8","unstructured":"Schuller, B., Rigoll, G., and Lang, M. (2004, January 17\u201321). Speech emotion recognition combining acoustic features and linguistic information in a hybrid support vector machine-belief network architecture. Proceedings of the 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing, Montreal, QC, Canada."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Gu, Y., Yang, K., Fu, S., Chen, S., Li, X., and Marsic, I. (2018, January 15\u201320). Multimodal affective analysis using hierarchical attention strategy with word-level alignment. Proceedings of the Conference Association for Computational Linguistics, Melbourne, Australia.","DOI":"10.18653\/v1\/P18-1207"},{"key":"ref_11","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Tsai, Y.-H.H., Bai, S., Liang, P.P., Kolter, J.Z., Morency, L.-P., and Salakhutdinov, R. (2019\u20132, January 28). Multimodal transformer for unaligned multimodal language sequences. Proceedings of the Conference Association for Computational Linguistics, Florence, Italy.","DOI":"10.18653\/v1\/P19-1656"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Huang, J., Tao, J., Liu, B., Lian, Z., and Niu, M. (2020, January 4\u20138). Multimodal transformer fusion for continuous emotion recognition. Proceedings of the ICASSP 2020\u20132020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053762"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Cho, K., Van Merri\u00ebnboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Han, W., Chen, H., Gelbukh, A., Zadeh, A., Morency, L.-P., and Poria, S. (2021, January 18\u201322). Bi-bimodal modality fusion for correlation-controlled multimodal sentiment analysis. Proceedings of the 2021 International Conference on Multimodal Interaction, Montreal, QC, Canada.","DOI":"10.1145\/3462244.3479919"},{"key":"ref_16","unstructured":"Hazarika, D., Zimmermann, R., and Poria, S. (2020, January 12\u201316). Misa: Modality-invariant and-specific representations for multimodal sentiment analysis. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Chiang, D., and Cholak, P. (2022, January 22\u201327). Overcoming a Theoretical Limitation of Self-Attention. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland.","DOI":"10.18653\/v1\/2022.acl-long.527"},{"key":"ref_18","unstructured":"Huang, Z., Xu, W., and Yu, K. (2015). Bidirectional LSTM-CRF models for sequence tagging. arXiv."},{"key":"ref_19","unstructured":"Zadeh, A.B., Liang, P.P., Poria, S., Cambria, E., and Morency, L.-P. (2018, January 15\u201320). Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Melbourne, Australia."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Hu, H., Xu, M.-X., and Wu, W. (2007, January 15\u201320). GMM supervector based SVM with spectral features for speech emotion recognition. Proceedings of the 2007 IEEE International Conference on Acoustics, Speech and Signal Processing-ICASSP\u201907, Honolulu, HI, USA.","DOI":"10.1109\/ICASSP.2007.366937"},{"key":"ref_21","unstructured":"Lin, Z., Feng, M., Santos, C.N.d., Yu, M., Xiang, B., Zhou, B., and Bengio, Y. (2017). A structured self-attentive sentence embedding. arXiv."},{"key":"ref_22","first-page":"34","article-title":"SVM scheme for speech emotion recognition using MFCC feature","volume":"69","author":"Milton","year":"2013","journal-title":"Int. J. Comput. Appl."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Han, K., Yu, D., and Tashev, I. (2014, January 7\u201310). Speech emotion recognition using deep neural network and extreme learning machine. Proceedings of the INTERSPEECH, Singapore.","DOI":"10.21437\/Interspeech.2014-57"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1016\/j.specom.2019.10.004","article-title":"Speech emotion recognition based on DNN-decision tree SVM model","volume":"115","author":"Sun","year":"2019","journal-title":"Speech Commun."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Badshah, A.M., Ahmad, J., Rahim, N., and Baik, S.W. (2017, January 13\u201315). Speech emotion recognition from spectrograms with deep convolutional neural network. Proceedings of the 2017 International Conference on Platform Technology and Service (PlatCon), Busan, Korea.","DOI":"10.1109\/PlatCon.2017.7883728"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Lee, J., and Tashev, I. (2015, January 6\u201310). High-level feature representation using recurrent neural network for speech emotion recognition. Proceedings of the INTERSPEECH, Dresden, Germany.","DOI":"10.21437\/Interspeech.2015-336"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Trigeorgis, G., Ringeval, F., Brueckner, R., Marchi, E., Nicolaou, M.A., Schuller, B., and Zafeiriou, S. (2016, January 20\u201325). Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472669"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Li, Y., Zhao, T., and Kawahara, T. (2019, January 15\u201319). Improved End-to-End Speech Emotion Recognition Using Self Attention Mechanism and Multitask Learning. Proceedings of the Interspeech 2019, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-2594"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wang, X., Wang, M., Qi, W., Su, W., Wang, X., and Zhou, H. (2021, January 6\u201311). A Novel end-to-end Speech Emotion Recognition Network with Stacked Transformer Layers. Proceedings of the ICASSP 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9414314"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Pang, B., Lee, L., and Vaithyanathan, S. (2002). Thumbs up? Sentiment classification using machine learning techniques. arXiv.","DOI":"10.3115\/1118693.1118704"},{"key":"ref_31","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Tang, D., Qin, B., and Liu, T. (2015, January 17\u201321). Document modeling with gated recurrent neural network for sentiment classification. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lisbon, Portugal.","DOI":"10.18653\/v1\/D15-1167"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Kim, Y. (2014). Convolutional Neural Networks for Sentence Classification. arXiv.","DOI":"10.3115\/v1\/D14-1181"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"e1253","DOI":"10.1002\/widm.1253","article-title":"Deep learning for sentiment analysis: A survey","volume":"8","author":"Zhang","year":"2018","journal-title":"WIREs Data Min. Knowl. Discov."},{"key":"ref_35","unstructured":"Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Huang, B., and Carley, K.M. (2019, January 3\u20137). Syntax-Aware Aspect Level Sentiment Classification with Graph Attention Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China.","DOI":"10.18653\/v1\/D19-1549"},{"key":"ref_37","unstructured":"Sun, C., Huang, L., and Qiu, X. (2019, January 2\u20137). Utilizing BERT for Aspect-Based Sentiment Analysis via Constructing Auxiliary Sentence. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, MN, USA."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lian, Z., Tao, J., Liu, B., Huang, J., Yang, Z., and Li, R. (2020, January 25\u201329). Conversational Emotion Recognition Using Self-Attention Mechanisms and Graph Neural Networks. Proceedings of the INTERSPEECH, Shangai, China.","DOI":"10.21437\/Interspeech.2020-1703"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Ghosal, D., Majumder, N., Poria, S., Chhaya, N., and Gelbukh, A. (2019). Dialoguegcn: A graph convolutional neural network for emotion recognition in conversation. arXiv.","DOI":"10.18653\/v1\/D19-1015"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Liu, J., Chen, S., Wang, L., Liu, Z., Fu, Y., Guo, L., and Dang, J. (2021, January 6\u201311). Multimodal emotion recognition with capsule graph convolutional based representation fusion. Proceedings of the ICASSP 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413608"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Peng, Z., Lu, Y., Pan, S., and Liu, Y. (2021, January 6\u201311). Efficient Speech Emotion Recognition Using Multi-Scale CNN and Attention. Proceedings of the ICASSP 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9414286"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Sun, L., Liu, B., Tao, J., and Lian, Z. (2021, January 6\u201311). Multimodal Cross-and Self-Attention Network for Speech Emotion Recognition. Proceedings of the ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9414654"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Cheng, J., Fostiropoulos, I., Boehm, B., and Soleymani, M. (2021, January 7\u201311). Multimodal Phased Transformer for Sentiment Analysis. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic.","DOI":"10.18653\/v1\/2021.emnlp-main.189"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Tang, J., Li, K., Jin, X., Cichocki, A., Zhao, Q., and Kong, W. (2021, January 1\u20136). CTFN: Hierarchical Learning for Multimodal Sentiment Analysis Using Coupled-Translation Fusion Network. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Bangkok, Thailand.","DOI":"10.18653\/v1\/2021.acl-long.412"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Wang, Z., Wan, Z., and Wan, X. (2020, January 20\u201324). Transmodality: An end2end fusion method with transformer for multimodal sentiment analysis. Proceedings of the Web Conference 2020, Taipei, Taiwan.","DOI":"10.1145\/3366423.3380000"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Hu, J., Liu, Y., Zhao, J., and Jin, Q. (2021). MMGCN: Multimodal fusion via deep graph convolution network for emotion recognition in conversation. arXiv.","DOI":"10.18653\/v1\/2021.acl-long.440"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Han, W., Chen, H., and Poria, S. (2021). Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multimodal Sentiment Analysis. arXiv.","DOI":"10.18653\/v1\/2021.emnlp-main.723"},{"key":"ref_48","unstructured":"Su, J. (2022, April 11). Entropy Invariance in Softmax Operation. Available online: https:\/\/kexue.fm\/archives\/9034."},{"key":"ref_49","unstructured":"Jang, E., Gu, S., and Poole, B. (2016). Categorical reparameterization with gumbel-softmax. arXiv."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"3878","DOI":"10.1121\/1.2935783","article-title":"Speaker identification on the SCOTUS corpus","volume":"123","author":"Yuan","year":"2008","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Degottex, G., Kane, J., Drugman, T., Raitio, T., and Scherer, S. (2014, January 4\u20139). COVAREP\u2014A collaborative voice analysis repository for speech technologies. Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853739"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Zadeh, A., Liang, P.P., Mazumder, N., Poria, S., Cambria, E., and Morency, L.-P. (2018, January 2\u20137). Memory fusion network for multi-view sequential learning. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12021"},{"key":"ref_53","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., and Antiga, L. (2019). Pytorch: An imperative style, high-performance deep learning library. arXiv."},{"key":"ref_54","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_55","unstructured":"Wang, Y., Shen, Y., Liu, Z., Liang, P.P., Zadeh, A., and Morency, L.-P. (February, January 27). Words can shift: Dynamically adjusting word representations using nonverbal behaviors. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_56","unstructured":"Pham, H., Liang, P.P., Manzini, T., Morency, L.-P., and P\u00f3czos, B. (February, January 27). Found in translation: Learning robust joint representations by cyclic translations between modalities. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_57","unstructured":"Iandola, F.N., Han, S., Moskewicz, M.W., Ashraf, K., Dally, W.J., and Keutzer, K. (2016). SqueezeNet: AlexNet-level accuracy with 50\u00d7 fewer parameters and <0.5 MB model size. arXiv."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201322). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_59","doi-asserted-by":"crossref","first-page":"149","DOI":"10.1613\/jair.731","article-title":"A model of inductive bias learning","volume":"12","author":"Baxter","year":"2000","journal-title":"J. Artif. Intell. Res."},{"key":"ref_60","unstructured":"Murphy, K.P. (2012). Machine Learning: A Probabilistic Perspective, MIT Press."},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Vu, T.-H., Jain, H., Bucher, M., Cord, M., and P\u00e9rez, P. (2019, January 15\u201320). Advent: Adversarial entropy minimization for domain adaptation in semantic segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00262"},{"key":"ref_62","unstructured":"Hjelm, R.D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y. (2018). Learning deep representations by mutual information estimation and maximization. arXiv."},{"key":"ref_63","unstructured":"Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. (2014). Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/7\/1010\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:55:55Z","timestamp":1760140555000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/7\/1010"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,21]]},"references-count":63,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2022,7]]}},"alternative-id":["e24071010"],"URL":"https:\/\/doi.org\/10.3390\/e24071010","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,21]]}}}