{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T15:07:45Z","timestamp":1773414465062,"version":"3.50.1"},"reference-count":39,"publisher":"World Scientific Pub Co Pte Ltd","issue":"01","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Patt. Recogn. Artif. Intell."],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:p>Multimodal Emotion Recognition (MER) aims to leverage information from multiple modalities to enhance the accuracy of emotion classification. However, existing methods struggle with properly disentangling shared and modality-specific features while maintaining discriminative power across modalities. To address this, we propose a novel contrastive learning-based MER framework that enforces inter-modal alignment while preserving intra-modal distinctiveness. Our method employs disentangled representation learning to extract both shared and private representations, ensuring that complementary features are effectively utilized. Additionally, we introduce an Adaptive Affinity Squeeze-Excitation (AASE) mechanism to dynamically refine multimodal representations by capturing inter-channel relationships, further enhancing the model\u2019s ability to identify crucial emotional cues. Experimental results demonstrate that our method improves emotion classification accuracy to 84.6% on benchmark datasets. Accordingly, the proposed MER framework provides a robust and highly adaptive solution for multimodal emotion analysis, effectively integrating modality-specific features while enhancing discriminability across modalities.<\/jats:p>","DOI":"10.1142\/s0218001425510231","type":"journal-article","created":{"date-parts":[[2025,11,19]],"date-time":"2025-11-19T05:45:37Z","timestamp":1763531137000},"source":"Crossref","is-referenced-by-count":1,"title":["Disentangled Representation and Contrastive Learning with Adaptive Affinity Squeeze-Excitation for Multimodal Emotion Recognition"],"prefix":"10.1142","volume":"40","author":[{"given":"Jhing-Fa","family":"Wang","sequence":"first","affiliation":[{"name":"Department of Electrical Engineering, National Cheng-Kung University, Tainan City 701401, Taiwan"},{"name":"School of Information Science and Technology, Sanda University, Shanghai City, P.\u00a0R.\u00a0China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7158-1725","authenticated-orcid":false,"given":"Hsin-Chun","family":"Tsai","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, Cheng Shiu University, Kaohsiung City 83347, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jie-Ming","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, National Cheng-Kung University, Tainan City 701401, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2025,11,19]]},"reference":[{"key":"S0218001425510231BIB001","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS60453.2023.00067"},{"key":"S0218001425510231BIB002","doi-asserted-by":"publisher","DOI":"10.1109\/FG.2019.8756532"},{"key":"S0218001425510231BIB003","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-008-9076-6"},{"key":"S0218001425510231BIB004","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01036"},{"key":"S0218001425510231BIB005","unstructured":"J. Devlin, BERT: Pre-training of deep bidirectional transformers for language understanding, preprint (2018), arXiv:1810.04805."},{"key":"S0218001425510231BIB006","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2024.112825"},{"key":"S0218001425510231BIB007","doi-asserted-by":"publisher","DOI":"10.1109\/ICCST53801.2021.00027"},{"key":"S0218001425510231BIB008","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351053"},{"key":"S0218001425510231BIB009","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413678"},{"key":"S0218001425510231BIB010","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2024.3356185"},{"key":"S0218001425510231BIB011","doi-asserted-by":"crossref","unstructured":"D. Hu, Y. Bao, L. Wei, W. Zhou and S. Hu, Supervised adversarial contrastive learning for emotion recognition in conversations, preprint (2023), arXiv:2306.01505.","DOI":"10.18653\/v1\/2023.acl-long.606"},{"key":"S0218001425510231BIB012","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-022-13091-9"},{"issue":"7","key":"S0218001425510231BIB013","first-page":"8419","volume":"45","author":"Lian Z.","year":"2023","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"S0218001425510231BIB014","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2022.11.076"},{"key":"S0218001425510231BIB015","doi-asserted-by":"crossref","unstructured":"Z. Liu, Y. Shen, V. B. Lakshminarasimhan, P. P. Liang, A. Zadeh and L.P. Morency, Efficient low-rank multimodal fusion with modality-specific factors, preprint (2018), arXiv:1806.00064.","DOI":"10.18653\/v1\/P18-1209"},{"key":"S0218001425510231BIB016","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1046"},{"issue":"1","key":"S0218001425510231BIB017","first-page":"164","volume":"34","author":"Mai S.","year":"2020","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"S0218001425510231BIB018","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3068598"},{"key":"S0218001425510231BIB019","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2022.3172360"},{"key":"S0218001425510231BIB020","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2014.2360798"},{"key":"S0218001425510231BIB021","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-023-15118-1"},{"key":"S0218001425510231BIB022","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.214"},{"key":"S0218001425510231BIB023","unstructured":"Y. Shou, T. Meng, W. Ai, N. Yin and K. Li, Adversarial representation with intra-modal and inter-modal graph contrastive learning for multimodal emotion recognition, preprint (2023), arXiv:2312.16778."},{"key":"S0218001425510231BIB024","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP48485.2024.10447667"},{"key":"S0218001425510231BIB025","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-022-14185-0"},{"key":"S0218001425510231BIB026","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1656"},{"key":"S0218001425510231BIB027","first-page":"1","volume":"30","author":"Vaswani A.","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"S0218001425510231BIB028","doi-asserted-by":"publisher","DOI":"10.1002\/int.22805"},{"issue":"1","key":"S0218001425510231BIB029","first-page":"7216","volume":"33","author":"Wang Y.","year":"2019","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"S0218001425510231BIB030","doi-asserted-by":"publisher","DOI":"10.1109\/ACII.2019.8925497"},{"key":"S0218001425510231BIB031","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547754"},{"key":"S0218001425510231BIB032","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3183587"},{"key":"S0218001425510231BIB033","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i12.17289"},{"key":"S0218001425510231BIB034","doi-asserted-by":"crossref","unstructured":"A. Zadeh, M. Chen, S. Poria, E. Cambria and L.P. Morency, Tensor fusion network for multimodal sentiment analysis, preprint (2017), arXiv:1707.07250.","DOI":"10.18653\/v1\/D17-1115"},{"issue":"1","key":"S0218001425510231BIB035","first-page":"5634","volume":"32","author":"Zadeh A.","year":"2018","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"S0218001425510231BIB036","first-page":"2236","volume-title":"Proc. 56th Annual Meeting of the Association for Computational Linguistics","volume":"1","author":"Zadeh A. B.","year":"2018"},{"key":"S0218001425510231BIB037","doi-asserted-by":"publisher","DOI":"10.1016\/j.cie.2022.108078"},{"key":"S0218001425510231BIB038","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2016.2598092"},{"key":"S0218001425510231BIB039","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2022.109978"}],"container-title":["International Journal of Pattern Recognition and Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218001425510231","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,15]],"date-time":"2025-12-15T05:35:19Z","timestamp":1765776919000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S0218001425510231"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,19]]},"references-count":39,"journal-issue":{"issue":"01","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["10.1142\/S0218001425510231"],"URL":"https:\/\/doi.org\/10.1142\/s0218001425510231","relation":{},"ISSN":["0218-0014","1793-6381"],"issn-type":[{"value":"0218-0014","type":"print"},{"value":"1793-6381","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,19]]},"article-number":"2551023"}}