{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T16:51:40Z","timestamp":1781283100964,"version":"3.54.1"},"reference-count":33,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2021,10,15]],"date-time":"2021-10-15T00:00:00Z","timestamp":1634256000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100007128","name":"Natural Science Foundation of Shaanxi Province","doi-asserted-by":"publisher","award":["2020JM-554"],"award-info":[{"award-number":["2020JM-554"]}],"id":[{"id":"10.13039\/501100007128","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Yan'an University Scientific Research Project","award":["YDY2019-18"],"award-info":[{"award-number":["YDY2019-18"]}]},{"DOI":"10.13039\/501100015401","name":"Key Research and Development Projects of Shaanxi Province","doi-asserted-by":"publisher","award":["2021NY-036"],"award-info":[{"award-number":["2021NY-036"]}],"id":[{"id":"10.13039\/501100015401","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Multi-modal fusion can achieve better predictions through the amalgamation of information from different modalities. To improve the performance of accuracy, a method based on Higher-order Orthogonal Iteration Decomposition and Projection (HOIDP) is proposed, in the fusion process, higher-order orthogonal iteration decomposition algorithm and factor matrix projection are used to remove redundant information duplicated inter-modal and produce fewer parameters with minimal information loss. The performance of the proposed method is verified by three different multi-modal datasets. The numerical results validate the accuracy of the performance of the proposed method having 0.4% to 4% improvement in sentiment analysis, 0.3% to 8% improvement in personality trait recognition, and 0.2% to 25% improvement in emotion recognition at three different multi-modal datasets compared with other 5 methods.<\/jats:p>","DOI":"10.3390\/e23101349","type":"journal-article","created":{"date-parts":[[2021,10,15]],"date-time":"2021-10-15T13:35:47Z","timestamp":1634304947000},"page":"1349","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["A Multi-Modal Fusion Method Based on Higher-Order Orthogonal Iteration Decomposition"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6099-3900","authenticated-orcid":false,"given":"Fen","family":"Liu\u00a0","sequence":"first","affiliation":[{"name":"School of Marine Science and Technology, Northwestern Polytechnical University, Xi\u2019an 710072, China"},{"name":"College of Mathematics and Computer Science, Yan\u2019an University, Yan\u2019an 716000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7281-7138","authenticated-orcid":false,"given":"Jianfeng","family":"Chen\u00a0","sequence":"additional","affiliation":[{"name":"School of Marine Science and Technology, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6590-5757","authenticated-orcid":false,"given":"Weijie","family":"Tan\u00a0","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Public Big Data, College of Computer Science and Technology, Guizhou University, Guiyang 550025, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chang","family":"Cai\u00a0","sequence":"additional","affiliation":[{"name":"School of Marine Science and Technology, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,15]]},"reference":[{"key":"ref_1","unstructured":"Fung, P.N. (2019, January 2). Modality-based Factorization for Multimodal Fusion. Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), Florence, Italy."},{"key":"ref_2","unstructured":"Shuang, W., Bondugula, S., Luisier, F., Zhuang, X., and Natarajan, P. (2014, January 23). Zero-Shot Event Detection Using Multi-modal Fusion of Weakly Supervised Concepts. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA."},{"key":"ref_3","first-page":"2013","article-title":"VideoStory Embeddings Recognize Events when Examples are Scarce","volume":"39","author":"Habibian","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1142\/S1793351X13400023","article-title":"Multimodal Information Fusion of Audio Emotion Recognition Based on Kernel Entropy Component Analysis","volume":"7","author":"Xie","year":"2013","journal-title":"Int. J. Semant. Comput."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Qi, J., and Peng, Y. (2018, January 13\u201319). Cross-modal Bidirectional Translation via Reinforcement Learning. Proceedings of the 27th International Joint Conference on Artificial Intelligence, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/365"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Bhagat, S., Uppal, S., Yin, Z., and Lim, N. (2020, January 23\u201328). Disentangling multiple features in video sequences using gaussian processes in variational autoencoders. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58592-1_7"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Tan, H., and Bansal, M. (2019). Lxmert: Learning cross-modality encoder representations from transformers. arXiv.","DOI":"10.18653\/v1\/D19-1514"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Yang, Z., Garcia, N., Chu, C., Otani, M., Nakashima, Y., and Takemura, H. (2020, January 1\u20135). Bert representations for video question answering. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Snowmass Village, CO, USA.","DOI":"10.1109\/WACV45572.2020.9093596"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Garcia, N., Otani, M., Chu, C., and Nakashima, Y. (2020, January 7\u20138). KnowIT VQA: Answering knowledge-based questions about videos. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6713"},{"key":"ref_10","first-page":"43:1","article-title":"A Review and Meta-Analysis of Multimodal Affect Detection Systems","volume":"47","author":"Kory","year":"2015","journal-title":"ACM Comput. Surv."},{"key":"ref_11","unstructured":"Kanluan, I., Grimm, M., and Kroschel, K. (2008, January 25\u201329). Audio-visual emotion recognition using an emotion space concept. Proceedings of the 16th European Signal Processing Conference, Lausanne, Switzerland."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Chetty, G., Wagner, M., and Goecke, R. (2015). A Multilevel Fusion Approach for Audiovisual Emotion Recognition, Wiley.","DOI":"10.1002\/9781118910566.ch17"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1109\/T-AFFC.2011.15","article-title":"DEAP: A Database for Emotion Analysis; Using Physiological Signals","volume":"3","author":"Koelstra","year":"2012","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1007\/s11042-013-1391-2","article-title":"Multimedia classification and event detection using double fusion","volume":"71","author":"Lan","year":"2014","journal-title":"Multimed. Tools Appl. Int. J."},{"key":"ref_15","first-page":"12136","article-title":"Deep multimodal multilinear fusion with high-order polynomial pooling","volume":"32","author":"Hou","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zadeh, A., Chen, M., Poria, S., Cambria, E., and Morency, L.P. (2017). Tensor Fusion Network for Multimodal Sentiment Analysis. arXiv.","DOI":"10.18653\/v1\/D17-1115"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Liu, Z., Shen, Y., Lakshminarasimhan, V.B., Liang, P.P., Zadeh, A., and Morency, L.P. (2018). Efficient Low-rank Multimodal Fusion with Modality-Specific Factors. arXiv.","DOI":"10.18653\/v1\/P18-1209"},{"key":"ref_18","first-page":"164","article-title":"Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal Fusion","volume":"34","author":"Mai","year":"2020","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"423","DOI":"10.1109\/TPAMI.2018.2798607","article-title":"Multimodal Machine Learning: A Survey and Taxonomy","volume":"41","author":"Baltrusaitis","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"63373","DOI":"10.1109\/ACCESS.2019.2916887","article-title":"Deep Multimodal Representation Learning: A Survey","volume":"7","author":"Guo","year":"2019","journal-title":"IEEE Access"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"455","DOI":"10.1137\/07070111X","article-title":"Tensor decompositions and applications","volume":"51","author":"Kolda","year":"2009","journal-title":"SIAM Rev."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Nojavanasghari, B., Gopinath, D., Koushik, J., BAltruaitis, T., and Morency, L.P. (2016, January 12\u201316). Deep Multimodal Fusion for Persuasiveness Prediction. Proceedings of the 18th ACM International Conference on Multimodal Interaction, Tokyo, Japan.","DOI":"10.1145\/2993148.2993176"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zadeh, A., Liang, P.P., Poria, S., Vij, P., and Morency, L.P. (2018, January 2\u20137). Multi-attention Recurrent Network for Human Communication Comprehension. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12024"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zadeh, A., Liang, P.P., Mazumder, N., Poria, S., and Morency, L.P. (2018, January 2\u20137). Memory Fusion Network for Multi-view Sequential Learning. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12021"},{"key":"ref_25","unstructured":"Zadeh, A., Zellers, R., Pincus, E., and Morency, L.P. (2016). MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Park, S., Han, S.S., Chatterjee, M., Sagae, K., and Morency, L.P. (2014, January 12\u201316). Computational Analysis of Persuasiveness in Social Multimedia: A Novel Dataset and Multimodal Prediction Approach. Proceedings of the 16th International Conference on Multimodal Interaction, Istanbul, Turkey.","DOI":"10.1145\/2663204.2663260"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: Interactive Emotional Dyadic Motion Capture Database","volume":"42","author":"Busso","year":"2008","journal-title":"Lang. Resour. Eval."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"3878","DOI":"10.1121\/1.2935783","article-title":"Speaker identification on the SCOTUS corpus","volume":"123","author":"Yuan","year":"2008","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Chen, M., Wang, S., Liang, P.P., Baltruaitis, T., Zadeh, A., and Morency, L.P. (2017, January 13\u201317). Multimodal Sentiment Analysis with Word-Level Fusion and Reinforcement Learning. Proceedings of the 19th ACM International Conference on Multimodal Interaction, Glasgow, UK.","DOI":"10.1145\/3136755.3136801"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C. (2014). Glove: Global Vectors for Word Representation. Proceeding of the 2014 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_31","unstructured":"Zhu, Q., Yeh, M.C., Cheng, K.T., and Avidan, S. (2006, January 17\u201322). Fast Human Detection Using a Cascade of Histograms of Oriented Gradients. Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, New York, USA."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"DeGottex, G., Kane, J., Drugman, T., Raitio, T., and Scherer, S. (2014, January 4\u20139). COVAREP: A Collaborative Voice Analysis Repository for Speech Technologies. Proceedings of the IEEE International Conference on Acoustics Speech and Signal Processing, Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853739"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/10\/1349\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:15:31Z","timestamp":1760166931000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/10\/1349"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,15]]},"references-count":33,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2021,10]]}},"alternative-id":["e23101349"],"URL":"https:\/\/doi.org\/10.3390\/e23101349","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,15]]}}}