{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T02:38:24Z","timestamp":1783478304331,"version":"3.55.0"},"reference-count":40,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2022,12,16]],"date-time":"2022-12-16T00:00:00Z","timestamp":1671148800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Natural Science Foundation of Shaanxi Province","award":["2020JM-554"],"award-info":[{"award-number":["2020JM-554"]}]},{"name":"Natural Science Foundation of Shaanxi Province","award":["YDY2019-18"],"award-info":[{"award-number":["YDY2019-18"]}]},{"name":"Yan\u2019an University Scientific Research Project","award":["2020JM-554"],"award-info":[{"award-number":["2020JM-554"]}]},{"name":"Yan\u2019an University Scientific Research Project","award":["YDY2019-18"],"award-info":[{"award-number":["YDY2019-18"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Multi-modal fusion can exploit complementary information from various modalities and improve the accuracy of prediction or classification tasks. In this paper, we propose a parallel, multi-modal, factorized, bilinear pooling method based on a semi-tensor product (STP) for information fusion in emotion recognition. Initially, we apply the STP to factorize a high-dimensional weight matrix into two low-rank factor matrices without dimension matching constraints. Next, we project the multi-modal features to the low-dimensional matrices and perform multiplication based on the STP to capture the rich interactions between the features. Finally, we utilize an STP-pooling method to reduce the dimensionality to get the final features. This method can achieve the information fusion between modalities of different scales and dimensions and avoids data redundancy due to dimension matching. Experimental verification of the proposed method on the emotion-recognition task using the IEMOCAP and CMU-MOSI datasets showed a significant reduction in storage space and recognition time. The results also validate that the proposed method improves the performance and reduces both the training time and the number of parameters.<\/jats:p>","DOI":"10.3390\/e24121836","type":"journal-article","created":{"date-parts":[[2022,12,19]],"date-time":"2022-12-19T05:55:43Z","timestamp":1671429343000},"page":"1836","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["A Parallel Multi-Modal Factorized Bilinear Pooling Fusion Method Based on the Semi-Tensor Product for Emotion Recognition"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6099-3900","authenticated-orcid":false,"given":"Fen","family":"Liu","sequence":"first","affiliation":[{"name":"School of Marine Science and Technology, Northwestern Polytechnical University, Xi\u2019an 710072, China"},{"name":"College of Mathematics and Computer Science, Yan\u2019an University, Yan\u2019an 716000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7281-7138","authenticated-orcid":false,"given":"Jianfeng","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Marine Science and Technology, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kemeng","family":"Li","sequence":"additional","affiliation":[{"name":"School of Marine Science and Technology, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6590-5757","authenticated-orcid":false,"given":"Weijie","family":"Tan","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Public Big Data, College of Computer Science and Technology, Guizhou University, Guiyang 550025, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chang","family":"Cai","sequence":"additional","affiliation":[{"name":"School of Marine Science and Technology, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Muhammad Saad","family":"Ayub","sequence":"additional","affiliation":[{"name":"School of Marine Science and Technology, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,12,16]]},"reference":[{"key":"ref_1","first-page":"423","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Ahuja","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","first-page":"2013","article-title":"VideoStory Embeddings Recognize Events when Examples are Scarce","volume":"39","author":"Habibian","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","unstructured":"Shuang, W., Bondugula, S., Luisier, F., Zhuang, X., and Natarajan, P. (2014, January 23\u201328). Zero-Shot Event Detection Using Multi-modal Fusion of Weakly Supervised Concepts. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Park, S., Han, S.S., Chatterjee, M., Sagae, K., and Morency, L.P. (2014, January 12\u201316). Computational Analysis of Persuasiveness in Social Multimedia: A Novel Dataset and Multimodal Prediction Approach. Proceedings of the 16th International Conference on Multimodal Interaction, New York, NY, USA.","DOI":"10.1145\/2663204.2663260"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zadeh, A., Chen, M., Poria, S., Cambria, E., and Morency, L.P. (2017). Tensor Fusion Network for Multimodal Sentiment Analysis. arXiv.","DOI":"10.18653\/v1\/D17-1115"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Liu, F., Chen, J.F., Tan, W.J., and Cai, C. (2021). A Multi-Modal Fusion Method Based on Higher-Order Orthogonal Iteration Decomposition. Entropy, 23.","DOI":"10.3390\/e23101349"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"8269683","DOI":"10.1155\/2020\/8269683","article-title":"The recognition of teacher behavior based on multimodal information fusion","volume":"2020","author":"Wu","year":"2020","journal-title":"Math. Probl. Eng."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Qi, J., and Peng, Y. (2018, January 13\u201319). Cross-modal Bidirectional Translation via Reinforcement Learning. Proceedings of the 27th International Joint Conference on Artificial Intelligence, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/365"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"3125879","DOI":"10.1155\/2018\/3125879","article-title":"Multimodal feature learning for video captioning","volume":"2018","author":"Lee","year":"2018","journal-title":"Math. Probl. Eng."},{"key":"ref_10","first-page":"1","article-title":"Multimodal Urban Sound Tagging with Spatiotemporal Context","volume":"2022","author":"Bai","year":"2022","journal-title":"IEEE Trans. Cogn. Dev. Syst."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"63373","DOI":"10.1109\/ACCESS.2019.2916887","article-title":"Deep Multimodal Representation Learning: A Survey","volume":"7","author":"Guo","year":"2019","journal-title":"IEEE Access"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1142\/S1793351X13400023","article-title":"Multimodal Information Fusion of Audio Emotion Recognition Based on Kernel Entropy Component Analysis","volume":"7","author":"Xie","year":"2013","journal-title":"Int. J. Semant. Comput."},{"key":"ref_13","unstructured":"Pang, L., and Ngo, C.W. Mutlimodal learning with deep boltzmann machine for emotion prediction in user generated videos. Proceedings of the 5th ACM on International Conference on Multimedia Retrieval."},{"key":"ref_14","unstructured":"Tsai, Y.H.H., Bai, S., Liang, P.P., Kolter, J.Z., Morency, L.P., and Salakhutdinov, R. Multimodal transformer for unaligned multimodal language sequences. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Sahay, S., Okur, E., Kumar, S.H., and Nachman, L. (2020). Low rank fusion based transformers for multimodal sequences. arXiv.","DOI":"10.18653\/v1\/2020.challengehml-1.4"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2617","DOI":"10.1109\/TASLP.2021.3096037","article-title":"Information fusion in attention networks using adaptive and multi-level factorized bilinear pooling for audio-visual emotion recognition","volume":"29","author":"Zhou","year":"2021","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"277","DOI":"10.1007\/s11042-009-0344-2","article-title":"Multimodal information fusion application to human emotion recognition from face and speech","volume":"49","author":"Mansoorizadeh","year":"2010","journal-title":"Multimed. Tools Appl."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"597","DOI":"10.1109\/TMM.2012.2189550","article-title":"Kernel cross-modal factor analysis for information fusion with application to bimodal emotion recognition","volume":"14","author":"Wang","year":"2012","journal-title":"IEEE Trans. Multimed."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Li, S., Zheng, W., Zong, Y., Lu, C., Tang, C., Jiang, X., and Xia, W. (2019, January 14\u201318). Bi-modality fusion for emotion recognition in the wild. Proceedings of the 19th International Conference on Multimodal Interaction, Suzhou, China.","DOI":"10.1145\/3340555.3355719"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Liu, C., Tang, T., Lv, K., and Wang, M. (2018, January 16\u201320). Multi-feature based emotion recognition for video clips. Proceedings of the 20th ACM International Conference on Multimodal Interaction, Boulder, CO, USA.","DOI":"10.1145\/3242969.3264989"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"570","DOI":"10.1109\/TMM.2008.921737","article-title":"Audio\u2013visual affective expression recognition through multistream fused HMM","volume":"10","author":"Zeng","year":"2008","journal-title":"IEEE Trans. Multimed."},{"key":"ref_22","unstructured":"Mai, S., Hu, H., and Xing, S. (2020, January 7\u201312). Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal Fusion. Proceedings of the 32th AAAI Conference on Artificial Intelligence, New York, NY, USA."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., and Rohrbach, M. (2016). Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding. arXiv.","DOI":"10.18653\/v1\/D16-1044"},{"key":"ref_24","unstructured":"Zadeh, A., Zellers, R., Pincus, E., and Morency, L.P. (2016). MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zadeh, A., Liang, P.P., Poria, S., Vij, P., and Morency, L.P. (2018, January 2\u20137). Multi-attention Recurrent Network for Human Communication Comprehension. Proceedings of the 32 AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12024"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, Z., Shen, Y., Lakshminarasimhan, V.B., Liang, P.P., Zadeh, A., and Morency, L.P. (2018). Efficient Low-rank Multimodal Fusion with Modality-Specific Factors. arXiv.","DOI":"10.18653\/v1\/P18-1209"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Yu, Z., Yu, J., Fan, J., and Tao, D. (2017, January 22\u201329). Multi-modal factorized bilinear pooling with co-attention learning for visual question answering. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.202"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zadeh, A., Liang, P.P., Mazumder, N., Poria, S., and Morency, L.P. (2018, January 2\u20137). Memory Fusion Network for Multi-view Sequential Learning. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12021"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1007\/BF02714570","article-title":"Semi-tensor product of matrices and its application to Morgen\u2019s problem","volume":"44","author":"Cheng","year":"2001","journal-title":"Sci. China Ser. Inf. Sci."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Fu, W., and Li, S. (2018, January 22\u201327). Semi-Tensor Compressed Sensing for Hyperspectral Image. Proceedings of the IEEE International Geoscience and Remote Sensing Symposium, Valencia, Spain.","DOI":"10.1109\/IGARSS.2018.8519360"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Bai, Z., Li, Y., Zhou, M., Li, D., Wang, D., Po\u0142ap, D., and Wo\u017aniak, M. (2020, January 19\u201324). Bilinear Semi-Tensor Product Attention (BSTPA) model for visual question answering. Proceedings of the 2020 International Joint Conference on Neural Networks, Glasgow, UK.","DOI":"10.1109\/IJCNN48605.2020.9206964"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1973","DOI":"10.1109\/TMM.2018.2794985","article-title":"A novel digital watermarking based on general non-negative matrix factorization","volume":"20","author":"Chen","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Cheng, D., Qi, H., and Zhao, Y. (2012). An Introduction to Semi-Tensor Product of Matrices and Its Applications, World Scientific.","DOI":"10.1142\/8323"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"2251","DOI":"10.1109\/TAC.2010.2043294","article-title":"A linear representation of dynamics of Boolean networks","volume":"55","author":"Cheng","year":"2010","journal-title":"IEEE Trans. Autom. Control"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1007\/BF02289464","article-title":"Some mathematical notes on three-mode factor analysis","volume":"31","author":"Tucker","year":"1966","journal-title":"Psychometrika"},{"key":"ref_36","first-page":"241","article-title":"Non-negative matrix factorization and its application in pattern recognition","volume":"51","author":"Liu","year":"2006","journal-title":"Chin. Sci. Bull."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1137\/S0036144598340483","article-title":"Two purposes for matrix factorization: A historical appraisal","volume":"42","author":"Hubert","year":"2000","journal-title":"SIAM Rev."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: Interactive Emotional Dyadic Motion Capture Database","volume":"42","author":"Busso","year":"2008","journal-title":"Lang. Resour. Eval."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C. (2014, January 25\u201329). Glove: Global Vectors for Word Representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"DeGottex, G., Kane, J., Drugman, T., Raitio, T., and Scherer, S. (2014, January 4\u20139). COVAREP: A Collaborative Voice Analysis Repository for Speech Technologies. Proceedings of the IEEE International Conference on Acoustics Speech and Signal Processing, Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853739"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/12\/1836\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:42:36Z","timestamp":1760146956000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/12\/1836"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,16]]},"references-count":40,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["e24121836"],"URL":"https:\/\/doi.org\/10.3390\/e24121836","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,16]]}}}